Information Processing System and Method Using Collective Intelligence
The information processing system leverages collective intelligence to label and process raw data and videos, addressing the inefficiencies in existing technologies by iteratively refining outputs through machine learning and additional labeling processes, thereby enhancing AI's learning and inference capabilities.
Patent Information
- Application Number
- JP2024565347
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-05-04
- Filing Date
- 2023-05-04
- Publication Date
- 2025-06-03
AI Technical Summary
Existing technologies lack an efficient method for labeling and processing low-level data and videos using collective intelligence, particularly in reconstructing operation-related videos of humans, avatars, and items into robot operation videos and enhancing the learning ability of artificial intelligence.
An information processing system and method that uses collective intelligence to label raw data, perform machine learning using classification and prediction models, and iteratively refine video outputs through additional labeling and learning processes.
The system effectively improves the inference and learning abilities of artificial intelligence by iteratively labeling and processing data and videos, enhancing the accuracy and efficiency of video generation and processing.
Smart Images

Figure 2025517145000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing system and method using collective intelligence, and particularly, performs labeling on one or more low-level data related to specific content provided by a user, performs a learning function on the labeled low-level data through a preset classification model and prediction model, performs additional labeling on a first video which is an output value of the prediction model, and performs an additional learning function on the additionally labeled first video through the classification model and the prediction model to output a second video. The present invention aims to provide an information processing system and method using collective intelligence.
Background Art
[0002] Collective intelligence is the intelligence obtained as a result of the intellectual ability accumulated by members of a group through cooperation or competition with each other, or indicates such a collective ability.
[0003] With the development of information database technologies such as avatars, items, and robotics, there is a need to link such collective intelligence with new big data-based knowledge services.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] An object of the present invention is to label one or more load data related to specific content provided by a user, perform a learning function on the labeled load data via a preset classification model and prediction model, perform additional labeling on a first video that is an output value of the prediction model, and perform an additional learning function via the classification model and the prediction model on the additionally labeled first video to output a second video, and to provide an information processing system and method using collective intelligence.
[0006] Another object of the present invention is to reconstruct operation-related videos of an actual person, a virtual avatar, an item, etc. as robot operation videos, label the reconstructed robot operation videos, perform a learning function on the labeled robot operation videos via a preset classification model and prediction model, perform additional labeling on a first robotics video that is a result of performing the learning function, and perform an additional learning function via the classification model and the prediction model on the additionally labeled first robotics video to output a second robotics video, and to provide an information processing system and method using collective intelligence.
Means for Solving the Problems
[0007] The information processing system using collective intelligence according to an embodiment of the present invention includes: a terminal that transmits one or more pieces of raw data collected in relation to a specific topic, meta information related to the raw data, a comparison target video, meta information related to the comparison target video, and identification information of a terminal; and a server that receives the one or more pieces of raw data related to a specific topic, meta information related to the raw data, a comparison target video, meta information related to the comparison target video, and identification information of a terminal transmitted from the terminal, performs selective labeling on the one or more pieces of raw data in conjunction with the terminal, performs machine learning on an artificial intelligence base based on information regarding the selectively labeled raw data, generates a classification value for the raw data based on the result of the machine learning, performs machine learning using, as input values, the classification value generated for the raw data, information regarding the selectively labeled raw data, the raw data, meta information related to the raw data, the comparison target video, and meta information related to the comparison target video, generates a first video corresponding to the raw data based on the result of the machine learning, and transmits the generated first video to the terminal.
[0008] As an example according to the present invention, the server may perform additional selective labeling on the first video in conjunction with the terminal, perform machine learning on an artificial intelligence base based on information regarding the additionally selectively labeled first video, generate a classification value for the first video based on the result of the machine learning, perform machine learning using, as input values, the classification value generated for the first video, information regarding the additionally selectively labeled first video, the first video, meta information related to the first video, the comparison target video, and meta information related to the comparison target video, generate a second video corresponding to the first video based on the result of the machine learning, and transmit the generated second video to the terminal.
[0009] As an example according to the present invention, the server may repeatedly perform the selection labeling execution process, the classification value generation process, the first video generation process, the additional selection labeling execution process for the generated first video, the additional classification value generation process for the generated first video, and the second video generation process on the plurality of raw data provided from a plurality of terminals in relation to the specific topic, and generate a second video that is collectively intelligent in relation to the specific topic.
[0010] The information processing method using collective intelligence according to an embodiment of the present invention includes steps of: receiving, by a server, one or more raw data related to a specific topic transmitted from a terminal, meta information related to the raw data, a comparison target video, meta information related to the comparison target video, and identification information of the terminal; performing, by the server, selection labeling on the one or more raw data in conjunction with the terminal; performing machine learning on an artificial intelligence base based on information related to the raw data that has been selection labeled, and generating a classification value for the raw data based on the result of the machine learning; performing machine learning by the server using, as input values, the classification value for the generated raw data, information related to the raw data that has been selection labeled, the raw data, meta information related to the raw data, the comparison target video, and meta information related to the comparison target video, and generating a first video corresponding to the raw data based on the result of the machine learning; transmitting, by the server, the generated first video to the terminal; and outputting, by the terminal, the first video transmitted from the server.
[0011] As an example according to the present invention, the step of performing selection labeling on the one or more raw data may set label values at at least one of one or more specific time points and one or more specific intervals of the raw data in response to a user input for the raw data displayed on the terminal.
[0012] As an example according to the present invention, the step of performing selective labeling on the one or more pieces of load data may be to set a label value for the correct or incorrect behavior of the movement of the object included in the load data at a specific time point or a specific section in response to the input of the user of the terminal with respect to the load data displayed in the video display area of the terminal.
[0013] As an example according to the present invention, before or after the step of performing selective labeling on the one or more pieces of load data by the server, the method may further include a step of performing hierarchical labeling on the one or more pieces of load data in conjunction with the terminal.
[0014] As an example according to the present invention, the step of performing hierarchical labeling on the one or more pieces of load data may include a process of setting a label value for at least one of another specific time point and another specific section of the load data in response to the input of the user based on a plurality of preset label classifications for the load data displayed on the terminal, and a process of dividing the load data into a plurality of sub-load data.
[0015] As an example according to the present invention, the step of generating a classification value for the load data based on the result of machine learning may perform machine learning using information related to the selectively labeled load data as an input value of a preset classification model, and generate a classification value for the load data based on the result of machine learning.
[0016] As an example according to the present invention, the step of generating a first video corresponding to the load data based on the result of machine learning may perform machine learning using the generated classification value for the load data, information related to the selectively labeled load data, the load data, meta information related to the load data, the comparison target video, and meta information related to the comparison target video as input values of a preset prediction model, and generate a first video related to the load data based on the result of machine learning.
[0017] As an example according to the present invention, the server performs additional selection labeling on the first video in conjunction with the terminal, and the server performs machine learning on the artificial intelligence base based on the information regarding the first video that has been additionally selected and labeled, and based on the result of the machine learning, generates a classification value for the first video. The server uses the generated classification value for the first video, the information regarding the first video that has been additionally selected and labeled, the first video, the meta information related to the first video, the comparison target video, and the meta information related to the comparison target video as input values to perform machine learning, and based on the result of the machine learning, generates a second video corresponding to the first video. The server transmits the generated second video to the terminal, the terminal outputs the second video transmitted from the server, and the server repeatedly performs the selection labeling process, the classification model inference process, the prediction model inference process, the additional selection labeling process for the generated first video, the additional classification model inference process, and the additional prediction model inference process respectively on the plurality of raw data provided from a plurality of terminals in relation to the specific topic, and generates a second video that has been collectively intellectualized in relation to the specific topic.
[0018] As an example according to the present invention, the step of performing additional selection labeling on the first video includes: the terminal divides the first video into a plurality of sub-videos based on information regarding sub-data that has been divided into a plurality by performing a hierarchical labeling function on the raw data; the terminal inputs label values for correct actions or label values for incorrect actions respectively for the plurality of divided sub-videos according to user input; the terminal inputs a label value indicating the order of the plurality of sub-videos according to user input in order to sort the order of the plurality of sub-videos; the terminal transmits the label values for correct and incorrect actions for the plurality of input sub-videos, the label value for sorting the order of the plurality of sub-videos, and the identification information of the terminal to the server; the server performs a time-series division selection labeling function on the first video and receives the label values for correct and incorrect actions for the plurality of sub-videos transmitted from the terminal, the label value for sorting the order of the plurality of sub-videos, and the identification information of the terminal. This may be included.
[0019] As an example according to the present invention, the step of performing additional selection labeling on the first video includes: a process in which the terminal divides the first video into a plurality of sub-videos based on information regarding sub-data divided into a plurality by performing a hierarchical labeling function on the raw data; a process in which the terminal inputs label values for the operation order of the avatars included in the plurality of divided sub-videos; a process in which the terminal inputs a label value indicating the order of the plurality of sub-videos in response to a user input in order to align the operation order by body part from the operations of the avatars included in the plurality of sub-videos; a process in which the terminal transmits the label values for the operation order of the avatars included in the plurality of input sub-videos, the label values for aligning the order of the plurality of sub-videos, and the identification information of the terminal to the server; and a process in which the server receives the label values for the operation order of the avatars included in the plurality of sub-videos transmitted from the terminal, the label values for aligning the order of the plurality of sub-videos, and the identification information of the terminal by performing a selection labeling function by body part on the first video. This may be included.
[0020] An information processing system using collective intelligence according to an embodiment of the present invention collects operation-related videos related to at least one of actual humans, avatars, and items in relation to a specific topic, and meta information related to the operation-related videos. To embody the collected operation-related videos as the operations of an actual robot, the collected operation-related videos are reconfigured as robot operation videos, and in conjunction with a terminal, selection labeling is performed on the robot operation videos. Machine learning of an artificial intelligence foundation is performed based on information regarding the selected-labeled robot operation videos, and based on the results of the machine learning, a classification value for the robot operation videos is generated. Based on the generated classification value for the robot operation videos, the information regarding the selected-labeled robot operation videos, the robot operation videos, the meta information related to the robot operation videos, comparison target videos, and the meta information related to the comparison target videos, a first robotics video corresponding to the robot operation videos is generated, and a server that transmits the generated first robotics video to the terminal, and the terminal that outputs the first robotics video transmitted from the server may be included.
[0021] As an example according to the present invention, the server performs additional selection labeling on the first robotics video in conjunction with the terminal, performs machine learning of an artificial intelligence foundation based on information regarding the additionally selected-labeled first robotics video, and based on the results of the machine learning, generates a classification value for the first robotics video. Machine learning is performed using, as input values, the generated classification value for the first robotics video, the information regarding the additionally selected-labeled first robotics video, the first robotics video, the meta information related to the first robotics video, comparison target videos, and the meta information related to the comparison target videos, and based on the results of the machine learning, a second robotics video corresponding to the first robotics video is generated, and the generated second robotics video may be transmitted to the terminal.
[0022] As an example according to the present invention, the server may repeatedly perform the selection labeling execution process, the classification value generation process, the first video generation process, the additional selection labeling execution process for the generated first robotics video, the additional classification inference value generation process for the generated first robotics video, and the second robotics video generation process for at least one of a plurality of actual humans, avatars, and items provided from a plurality of terminals in relation to the specific topic, and generate a second robotics video that is collectively intelligent in relation to the specific topic.
[0023] An information processing method using collective intelligence according to an embodiment of the present invention includes steps of: collecting, by a server, an action-related video related to at least one of an actual human, an avatar, and an item, and meta information related to the action-related video, in relation to a specific topic; reconstructing, by the server, the collected action-related video as a robot action video in order to embody the collected action-related video as an action of an actual robot; performing, by the server in conjunction with a terminal, selection labeling on the robot action video; performing machine learning on an artificial intelligence base based on information about the robot action video subjected to the selection labeling, and generating a classification value for the robot action video based on a result of the machine learning; generating, by the server, a first robotics video corresponding to the robot action video based on the classification value for the generated robot action video, the information about the robot action video subjected to the selection labeling, the robot action video, the meta information related to the robot action video, a comparison target video, and meta information related to the comparison target video; transmitting, by the server, the generated first robotics video to the terminal; and outputting, by the terminal, the first robotics video transmitted from the server.
[0024] As an example according to the present invention, before or after the step of performing selective labeling on the robot operation video by the server, a step of performing hierarchical labeling on the robot operation video in conjunction with the terminal may be further included.
[0025] As an example according to the present invention, a step of performing additional selective labeling on the first robotics video in conjunction with the terminal by the server; a step of performing machine learning of an artificial intelligence base based on information regarding the first robotics video that has been additionally selectively labeled, and generating a classification value for the first robotics video based on the result of the machine learning; a step of performing machine learning with the classification value generated for the first robotics video, the information regarding the first robotics video that has been additionally selectively labeled, the first robotics video, the meta information related to the first robotics video, the comparison target video, and the meta information related to the comparison target video as input values, and generating a second robotics video corresponding to the first robotics video based on the result of the machine learning; a step of transmitting the generated second robotics video to the terminal by the server; a step of outputting the second robotics video transmitted from the server by the terminal; and a step of repeatedly performing the selective labeling execution process, the classification value generation process, the first robotics video generation process, the additional selective labeling execution process for the generated first robotics video, the additional classification value generation process for the generated first robotics video, and the first robotics video generation process for at least one of a plurality of actual humans, avatars, and items provided from a plurality of terminals in relation to an operation-related video associated with a specific topic, and generating a second robotics video that has been collectively intelligentized in relation to the specific topic may be further included.
Effects of the Invention
[0026] The present invention performs labeling on one or more pieces of raw data related to specific content provided by a user, performs a learning function on the labeled raw data via a preset classification model and prediction model, performs additional labeling on a first video that is an output value of the prediction model, and performs an additional learning function via the classification model and prediction model on the additionally labeled first video to output a second video, thereby providing an avatar and / or item related to the raw data to the user, and having an effect of improving the inference ability of artificial intelligence by labeling the raw data.
[0027] In addition, the present invention reconstructs operation-related videos of actual humans, virtual avatars, items, etc. as robot operation videos, performs labeling on the reconstructed robot operation videos, performs a learning function on the labeled robot operation videos via a preset classification model and prediction model, performs additional labeling on a first robotics video that is a result of performing the learning function, and performs an additional learning function via the classification model and prediction model on the additionally labeled first robotics video to output a second robotics video, thereby having an effect of improving the learning ability of artificial intelligence by repeatedly applying the result of artificial intelligence to the classification model and prediction model of artificial intelligence.
Brief Description of the Drawings
[0028]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32
Embodiments for Carrying Out the Invention
[0029] It should be noted that the technical terms used in the present invention are merely used to explain specific embodiments and are not intended to limit the present invention. Further, unless otherwise defined in the present invention, the technical terms used in the present invention should be interpreted in the meaning generally understood by those with ordinary knowledge in the technical field to which the present invention pertains, and should not be interpreted in an overly comprehensive meaning or an overly narrowed meaning. Also, when the technical terms used in the present invention are incorrect technical terms that cannot accurately express the idea of the present invention, they should be understood as being replaced by technical terms that can be correctly understood by those skilled in the art. Further, general terms used in the present invention should be analyzed according to the content defined in the dictionary or according to the context, and should not be interpreted in an overly narrowed meaning.
[0030] Also, the singular expressions used in the present invention include plural expressions unless the context clearly indicates a different meaning. Terms such as "composed of" or "including" in the present invention should not necessarily be analyzed as including all of the various components or many steps described in the invention. Some of the components or some of the steps may not be included, or it should be analyzed that additional components or steps may be further included.
[0031] Also, terms including ordinal numbers such as first, second, etc. used in the present invention may be used to explain components, but the components should not be limited by the terms. The terms are only used for the purpose of distinguishing one component from another. For example, without departing from the scope of the rights of the present invention, the first component may be named the second component, and similarly, the second component may also be named the first component.
[0032] Hereinafter, preferred embodiments according to the present invention will be described in detail with reference to the accompanying drawings. However, the same or similar components are given the same reference numerals regardless of the reference signs, and duplicate explanations thereof are omitted.
[0033] In addition, when it is determined that a specific description of such known technology may obscure the gist of the present invention in explaining the present invention, the detailed description thereof will be omitted. It should also be noted that the accompanying drawings are only for facilitating the understanding of the idea of the present invention, and the idea of the present invention should not be construed as being limited by the accompanying drawings.
[0034] FIG. 1 is a block diagram showing the configuration of an information processing system 10 using collective intelligence according to an embodiment of the present invention.
[0035] As shown in FIG. 1, the information processing system 10 using collective intelligence is composed of a terminal 100 and a server 200. Not all of the components of the information processing system 10 using collective intelligence shown in FIG. 1 are essential components, and the information processing system 10 using collective intelligence may be embodied by more components than those shown in FIG. 1, or may also be embodied by fewer components.
[0036] The terminal 100 may be applied to various terminals such as a smart phone, a portable terminal, a mobile terminal, a foldable terminal, a personal digital assistant (PDA (registered trademark)), a PMP (Portable Multimedia Player) terminal, a telematics terminal, a navigation terminal, a personal computer, a notebook computer, a slate PC, a tablet PC, an ultrabook, a wearable device (including, for example, a smartwatch, a smart glass, an HMD (Head Mounted Display), etc.), a wibro terminal, an IPTV (Internet Protocol Television) terminal, a smart TV, a digital broadcast terminal, an AVN (Audio Video Navigation) terminal, an A / V (Audio / Video) system, a flexible terminal, a digital signage device, a VR simulator, a robot, etc.).
[0037] The server 200 may be embodied in the form of cloud computing, grid computing, server-based computing, utility computing, network computing, quantum cloud computing, a web server, a database server, a proxy server, etc. Further, one or more of a network load balancing mechanism or various software that enables the corresponding server 200 to operate on the Internet or other networks may be installed in the server 200, and it may be embodied as a computerized system through this. Also, the network may be an http network, a private line, an intranet, or any other network. Furthermore, the connection between the terminal 100 and the server 200 may be connected through a security network so that data is not subject to attacks by any hacker or other third party. Also, the server 200 may include a plurality of database servers, and such database servers may be embodied in a manner of being separately connected to the server 200 via any type of network connection including the architecture of a distributed database server.
[0038] Each of the terminal 100 and the server 200 may include a communication unit (not shown) for performing communication functions with other terminals, a storage unit (not shown) for storing various information and programs (or applications), a display unit (not shown) for displaying various information and program execution results, an audio output unit (not shown) for outputting audio information corresponding to the various information and program execution results, a control unit (not shown) for controlling various components and functions of each terminal, etc.
[0039] The terminal 100 communicates with the server 200 and the like. At this time, the terminal 100 may be a terminal owned by a user (or an expert in a specific field) for performing functions such as a raw data collection function, a hierarchical labeling function for information / videos, a selective labeling function for information / videos, a time-series segmentation selective labeling function for information / videos, and a selective labeling function for body parts of information / videos via a dedicated application provided by the corresponding server 200.
[0040] In addition, the terminal 100 registers as a member as a user for providing functions such as a raw data collection function, a hierarchical labeling function for information / videos, a selective labeling function for information / videos, a time-series segmentation selective labeling function for information / videos, and a selective labeling function for body parts of information / videos via a dedicated application and / or website provided by the server 200 in cooperation with the server 200, and registers personal information and the like with the server 200. At this time, the personal information includes an ID, an email address, a password (or PIN), a name, a gender, a date of birth, a contact information, a residential address (or address information), and the like.
[0041] In addition, the terminal 100 can also register as a member as a user with the server 200 by using SNS account information, account information of other websites, or account information of mobile messengers registered by the user of the corresponding terminal 100. Here, the SNS account may be information related to Facebook, Twitter (registered trademark), Instagram, Kakao Story, Naver Blog, and the like. Also, the account of the other website may be information related to YouTube, Kakao, Naver, and the like. Also, the mobile messenger account may be information related to KakaoTalk, Line, Viber, WeChat (wechat (registered trademark)), WhatsApp, Telegram, Snapchat, and the like.
[0042] Also, when performing the membership registration procedure, the terminal 100 can successfully complete the membership registration procedure to the server 200 only after completing the authentication function through the personal authentication means (including, for example, mobile phones, credit cards, IPINs, etc.).
[0043] Also, after the membership registration is completed, the terminal 100 installs the dedicated app (or application / application program / specific app) provided by the server 200 on the corresponding terminal 100 in order to use the services provided by the server 200. At this time, the dedicated app may include native apps, mobile web apps, responsive web apps, adaptive web apps, hybrid apps, etc., and may be an app for performing functions such as load data collection functions, hierarchical labeling functions for information / videos, selection labeling functions for information / videos, time-series segmentation selection labeling functions for information / videos, and selection labeling functions for different body parts of information / videos.
[0044] Also, after the membership registration is completed, the terminal 100 can display the discount coupons provided by the server 200 through the corresponding dedicated app. At this time, the discount coupons may be discount coupons including a certain ratio of discount information when using functions such as the load data collection function, hierarchical labeling function for information / videos, selection labeling function for information / videos, time-series segmentation selection labeling function for information / videos, and selection labeling function for different body parts of information / videos provided by the corresponding server 200.
[0045] In addition, in order to perform the functions provided by the server 200, the terminal 100 performs a settlement function in conjunction with the server 200 and a settlement server (not shown) according to the subscription function. At this time, the server 200 can perform the settlement function by means of card settlement, automatic transfer through linkage with a bank settlement account, settlement using the cash points or cash remaining in the account of the terminal 100 registered as a member with the server 200, simple settlement including Kakao Pay, Naver Pay, etc.
[0046] When the settlement fails, the terminal 100 receives information (including, for example, insufficient balance, limit exceeded, etc.) indicating that the settlement has failed, transmitted from the server 200 (or the settlement server), and outputs (or displays) the received information indicating that the settlement has failed.
[0047] In addition, after the settlement function is normally performed, the terminal 100 receives the execution result of the settlement function transmitted from the server 200. Here, the execution result of the settlement function includes the subscription period, settlement amount, settlement date and time information, etc.
[0048] In addition, the terminal 100 executes a dedicated application pre-installed on the corresponding terminal 100 and displays an application execution result screen obtained by executing the dedicated application. Here, the application execution result screen includes a collection menu (or button / item) for collecting one or more pieces of raw data related to a specific topic, meta information related to the corresponding raw data, etc., a view menu for displaying the collected information and the information provided by the server 200, a setting menu for environment settings, and the like. Here, the terminal 100 performs a login procedure when executing the dedicated application by using the ID and password obtained by membership registration, a barcode or QR code (registered trademark) including the ID, etc. to a server 200 that provides the corresponding dedicated application in a state of having registered as a member, and can perform one or more functions of the corresponding dedicated application (for example, including a raw data collection function, a hierarchical labeling function for information / videos, a selection labeling function for information / videos, a time-series division selection labeling function for information / videos, a selection labeling function for body parts for information / videos, etc.).
[0049] In addition, when a pre-set collection menu on the application execution result screen displayed on the terminal 100 is selected, the terminal 100 displays a collection screen corresponding to the selected collection menu in order to collect one or more pieces of raw data related to a specific topic, meta information related to the corresponding raw data, a comparison target video, meta information related to the corresponding comparison target video, etc. from one or more visual set devices (not shown) according to the user's settings. Here, the collection screen includes a selection item for an information collection target for selecting one or more visual set devices linked to the corresponding terminal 100 according to the user's selection (or the user's input / touch / control), a selection item for the type of collection information for selecting the type of information to be collected from the selected information collection target, a collection start item for collecting information from the selected item according to the selected information collection target, and the like.
[0050] In addition, the terminal 100 receives a plurality of input values corresponding to a plurality of input items in response to an input (or selection / touch / control by the user / expert) of the user of the corresponding terminal 100 on the collection screen displayed on the corresponding terminal 100. Here, the plurality of input values include an information collection target (or visual set device information / identification information of the visual set device), the type of information to be collected (for example, including sequential still images (or a plurality of sequential still images), videos, measurement values / sensor values, etc.).
[0051] In addition, based on the plurality of received input values, the terminal 100 cooperates with the one or more visual set devices and collects one or more load data, meta information related to the corresponding load data, a comparison target video, meta information related to the corresponding comparison target video, etc. in relation to a specific topic. Here, the specific topic (or specific content) includes medical acts (for example, including surgeries, operations, etc.), dances, sports events (for example, including soccer, basketball, table tennis, etc.), games, e-sports, etc. At this time, the terminal 100 can also collect one load data from one user in relation to the corresponding specific topic, or can collect a plurality of different load data (or annotation step or attribute item load data / basic video information) from one user. Here, the comparison target video may be content that does not conflict with intellectual property rights such as copyright and portrait rights.
[0052] The visual set device communicates with the terminal 100, the server 200, etc.
[0053] In addition, the visual set device includes a camera unit, a lidar, an eye tracker, a motion capture and a motion tracker, medical equipment (for example, CT, scanner, MRI, medical ultrasound, etc.).
[0054] In addition, the visual set device acquires (or collects / shoots / measures) actual real-world images (or actual real-world image information) related to the location (or area) where the corresponding visual set device is configured (or arranged / installed). Here, the actual real-world images indicate raw data (or original data / source data / visual data), and include sequential still images (or a plurality of sequential still images / attributes), videos (or target attributes), measurement values, etc. that are acquired (or collected / shooted / measured) in the actual real world. Also, the measurement values include video information (or 3D data) measured by the lidar, the eye tracker, the motion capture and motion tracker, the medical equipment, etc. Further, the one or more acquired actual real-world images can be merged and used.
[0055] In addition, the terminal 100 can also acquire actual real-world images in conjunction with Cinematic Reality, which is a medical assistance application developed by Siemens Healthineers using the Microsoft HoloLens 2. Here, Cinematic Reality includes the function of rendering voxel data obtained from medical CT, MRI, etc. The data rendered by Cinematic Reality is used as a dataset for manufacturing digital casts, 3D printing artificial casts, etc. At this time, the voxel data is used in combination with point cloud data in the form of GNN.
[0056] FIG. 2 is a diagram showing the raw data according to an embodiment of the present invention. Here, the raw data includes actual real-world images (or actual real-world data), robot operation images (or robot operation image information), etc. At this time, the robot operation images are collected by the visual set device during the actual operation of the robot, and are applied to FIGS. 1 to 17 and FIG. 22 in the same manner as the raw data of the avatar and / or item.
[0057] The raw data shown in FIG. 2 above represents K1 clusters (or sequential data / still images). Here, K may be a natural number (or a positive integer). At this time, the virtual generated data (Augmentation data) generated by the server 200 is included in the corresponding raw data.
[0058] In addition, the virtual generated data generated using the raw data in FIG. 2 is provided as an attribute item (or multiple data in the annotation step).
[0059] The primary objective of the present invention in virtual surgery simulation and virtual tooth extraction simulation is to maximize performance using a small amount of actual surgical collection data (or actual real - world video information / raw data). For this purpose, the virtual digital cadaver generated data is provided in the training and simulation steps, and the artificial intelligence (or classification model / prediction model) can be trained with teacher supervision in a way that the doctor selects and labels the virtual digital cadaver generated data.
[0060] The digital cadaver is a patient avatar. To compensate for the difficulty of reflecting the different body structures (or variations) of individual patients in the uniform digital attributes, which is a shortcoming of the digital cadaver, medical information collected from information - gathering devices at the medical site (or the visual set device) (for example, including CT, X - Ray, ultrasonic devices, oral scanners, etc.), the experience of professionals, and the knowledge of professionals are utilized. By using such comprehensive information and concurrently using digital cadavers and artificial cadavers that reflect the variations of specific patients, virtual reality (VR) and virtual treatments and virtual surgeries of 3D simulators (not shown) can be advanced.
[0061] Further, the terminal 100 transmits to the server 200 one or more pieces of load data related to the collected specific topic, meta information related to the corresponding load data, a comparison target video, meta information related to the corresponding comparison target video, identification information of the terminal 100, and the like. Here, the identification information of the terminal 100 includes MDN (Mobile Directory Number), mobile IP, mobile MAC, Sim (subscriber identity module) card unique information, serial number, and the like.
[0062] At this time, when the comparison target video related to the load data is not collected by the corresponding terminal 100, the terminal 100 transmits to the server 200 one or more pieces of load data related to the collected specific topic, meta information related to the corresponding load data, identification information of the terminal 100, and the like.
[0063] Further, the terminal 100 receives from the server 200 a comparison target video related to the corresponding load data transmitted in response to the transmission, meta information related to the corresponding comparison target video, and the like, and matches (or maps / links) and manages the received comparison target video related to the corresponding load data, meta information related to the corresponding comparison target video, and the like with one or more pieces of load data related to the collected specific topic, meta information related to the corresponding load data, and the like.
[0064] Further, the terminal 100 displays (or outputs) one or more pieces of load data related to the collected specific topic, meta information related to the corresponding load data, a comparison target video, meta information related to the corresponding comparison target video, and the like. At this time, the terminal 100 can also apply virtual reality, augmented reality, extended reality, mixed reality, etc. to the corresponding load data and the like for display (or output).
[0065] That is, when a view menu preset on the execution result screen of the application displayed on the terminal 100 is selected, the terminal 100 displays a view screen corresponding to the selected view menu in order to display the collected information and the information provided by the server 200. Here, the view screen includes a video display area for displaying the raw data and the generated video, a display area for the comparison target video for displaying the comparison target video, a hierarchical label input menu for selecting variable values (or label values) for hierarchical labeling, a selection label input menu for selecting set values for selection labeling, a playback bar for providing functions such as play / pause / stop for the video, and the like.
[0066] Also, when the playback bar included in the view screen within the execution result screen of the application displayed on the terminal 100 is selected, or when the playback button within the corresponding view screen is selected, the terminal 100 displays (or outputs) the collected raw data in the video display area, and displays (or outputs) the comparison target video corresponding to the collected raw data (or the comparison target video corresponding to the corresponding raw data provided by the server 200) in the display area for the comparison target video. At this time, the terminal 100 may perform synchronization for the corresponding raw data and the comparison target video based on the meta information corresponding to the raw data and the comparison target video respectively, and display the synchronized raw data and comparison target video in the video display area and the display area for the comparison target video respectively. Here, when either one of the raw data displayed in the video display area within the terminal 100 and the comparison target video displayed in the display area for the comparison target video stops due to a pause function or a stop function, the terminal 100 controls the other one to also stop by the pause function or the stop function.
[0067] Also, the terminal 100, in conjunction with the server 200, sets (or receives / inputs) a label (or label value) at a specific time point (or specific section) of the corresponding raw data in response to an input (or user selection / touch / control) of the user of the corresponding terminal 100 for the raw data displayed on the corresponding terminal 100.
[0068] Further, for the load data displayed in the video display area of the terminal 100, the terminal 100 sets (or receives / inputs) a label (or label value) for the correct action or incorrect action with respect to the movement (or action of the object) of the object included in the corresponding load data at a specific time point (or specific section) in response to the input of the user of the corresponding terminal 100 (or the user's selection / touch / control).
[0069] That is, at one or more specific time points of the load data displayed in the video display area, the terminal 100 inputs a label value for the correct action (for example, a preset approval / consent / ACCEPT label) or a label value for the incorrect action (for example, a preset rejection / REJECT label) in response to the user's input respectively.
[0070] In this way, for the load data related to the corresponding specific topic, the terminal 100 sets (or receives / inputs) one or more selection labels (or selection label values) at one or more specific time points (or specific sections) in response to the input of the user of the corresponding terminal 100 who is an expert related to the corresponding specific topic.
[0071] Also, in this way, the user of the terminal 100 judges based on his own expertise from the load data displayed (or output) on the corresponding terminal 100. When a part related to an incorrect action is visible, the user selects a rejection label for the corresponding part and selects an approval label for the part related to the correct action.
[0072] In addition, the terminal 100 can label, in a binary, ternary, or multi-way manner, a specific point (or specific interval) of the load data displayed on the corresponding terminal 100 by using an object recognition method that drags (drag) a mouse (not shown) or attaches a tag to the corresponding load data and automatically recognizes the boundary line and boundary surface. Here, the selection labeling (or selection re-labeling / primary selection labeling / first selection labeling) refers to a labeling method for setting (or attaching) a label (or label value) for the presence or absence of an error (or abnormality) at a specific point (or specific interval) of the load data. At this time, for a point (or interval) of the load data for which a label (or label value) has not been set by the selection labeling, a preset default label value (for example, an approval label) may be set. In addition, the terminal 100 can attach a preset not ACCEPT label to a point (or interval / attribute / target attribute) of the load data for which the approval label is not attached, and can also attach a preset not REJECT label to a point (or interval / attribute / target attribute) of the load data for which the rejection label is not attached.
[0073] In addition, for object recognition (object recognition / object detection), when a user of the corresponding terminal 100 drags or attaches a tag to the load data displayed on the corresponding terminal 100, the artificial neural network detects one or more incorrect operation sites and operations, separates and analyzes the image. In addition, the terminal 100 provides the inference result to the user through the artificial intelligence inference process.
[0074] In various embodiments, the load data (or video information) includes 2D video information, 3D video information, point cloud information of a still image, and the like.
[0075] In addition, the terminal 100 moves the cursor (or mouse point) of the mouse in accordance with the user's input on the timeline within the corresponding playback bar using the raw data (e.g., video) displayed on the corresponding terminal 100, pauses at a specific point in time, captures a still image, and then automatically recognizes the boundary line and boundary surface in the still image at the corresponding time and attaches tags using the mouse button and cursor. Further, when the terminal 100 attaches tags to a plurality of 3D still images captured from a video, it controls so that the boundary surface and boundary line of the entire video are automatically recognized.
[0076] Also, when it is desired to attach an approval label or a rejection label to the entire raw data (or video information) output to the terminal 100, the terminal 100 can be operated by directly pressing an approval button or a rejection button displayed on the corresponding terminal 100 in accordance with the user's input. When it is desired to attach a label by pressing an approval button or a rejection button after specifying a fine part, a boundary line (including, for example, a straight line, a curve, etc.) or a boundary surface (including, for example, a closed curve, etc.) can be specified using mouse dragging, or after specifying a plurality of points with the mouse button, an approval button or a rejection button can be pressed to attach a label.
[0077] In one embodiment of the present invention, in the object recognition method, object detection, position measurement, object and instance segmentation, pose estimation, etc. are applied, and are similarly applied to instance tracking, action recognition, motion estimation, etc. for video analysis. Further, it is used in combination with a convolutional neural network to sense the actions included in the video clip. Action sensing, scene extraction, next frame prediction, object tracking, etc. are used. Based on the automatically recognized boundary line and boundary surface, the corresponding label is attached by pressing an approval button or a rejection button to the correct part and the incorrect part of the object and action output from the interface, respectively.
[0078] In one embodiment of the present invention, when the left button of the mouse is pressed to tag and / or dragged to tag a plurality of point clouds of 2D video information and 3D video information, the boundary lines and boundary surfaces may be automatically recognized. Also, when the left button of the mouse is pressed to tag or dragged to tag a plurality of point clouds existing on the x, y, and z coordinates among the 3D still images, the boundary line between correct information and incorrect information is automatically recognized, and the boundary surface can be automatically recognized by a closed curve.
[0079] Further, the information processing system 10 may further include other input devices (not shown).
[0080] The other input devices communicate with the terminal 100, the server 200, and the like.
[0081] Further, the other input devices are used when tagging or dragging to label the load data (or video information).
[0082] Further, the other input devices include a controller, an eye tracker, a data glove, a speech recognition interface, a Brain-Computer Interface (BCI), a hand tracking technology, a haptic device, and the like.
[0083] Next, an example of the usage method of the other input devices is shown.
[0084] That is, the method of using the other input device includes a method of operating the cursor and buttons of a mouse with a voice recognition interface and a brain-computer interface to attach tags or drag and label, a method of operating a controller that emits light rays with a voice recognition interface and a brain-computer interface to attach tags or drag and label, a method of operating an eye tracker with a voice recognition interface and a brain-computer interface to attach tags or drag and label, a method of operating a data glove with hand tracking technology, a voice recognition interface, and a brain-computer interface to attach tags or drag and label, and the like.
[0085] In one embodiment of the present invention, the buttons of a computer mouse are directly operated with a voice recognition interface. It is also possible to move the light rays of the controller to attach tags or drag. In addition, an eye tracker is used to sense the user's line of sight, and the object to be classified is identified by attaching a tag to the center of the viewing angle, and the object is labeled. Further, by using the interaction between the data glove and hand movements (or hand movement tracking technology), tags are attached to the video on the user interface, or boundary lines and boundary surfaces are created for the object. When a platform user (or a group of experts in each field) uses the brain-computer interface technology that connects the human brain and the computer, the platform user can view a video (including, for example, still images, videos, etc.), attach tags or drag with a mouse to the boundary lines and boundary surfaces, and then label by pressing an approval button or a rejection button on the still image or video only with their own will (or thought). Even further, after attaching tags or dragging for creating boundary lines and boundary surfaces in the still image or video only with their own will (or thought), selective labeling is performed on the classified still images and videos.
[0086] In one embodiment of the present invention, when the brain-computer interface is integrated with recurrent neural networks, convolutional neural networks, multi-layer neural network algorithms, and robot arm technology, just by thinking, one can press the approval button or rejection button displayed on the screen of the corresponding terminal 100 to attach a label or perform selective labeling. Further, the terminal 100 can also label still image information and video information using a brain-machine interface, a neuromorphic chip, etc., and perform hierarchical clustering based on the labeled information.
[0087] In one embodiment of the present invention, when an advanced brain-computer interface is developed, the user interface (or screen) displayed on the corresponding terminal 100 can appear in the user's mind just by thinking, and labeling can also be performed just by the user's thought. Further, the terminal 100 or the server 200 can perform hierarchical clustering based on the labeled information and utilize it for classification models and prediction models.
[0088] In one embodiment of the present invention, the method of designating incorrect parts in a still image is as follows.
[0089] When a dentist determines that the position of the orthodontic mini-implant implanted in a patient's oral cavity is slightly higher or lower than the appropriate position based on their medical knowledge, they can use mouse dragging to specify a boundary line (including, for example, a straight line, a curve, etc.) or a boundary surface (including, for example, a closed curve, etc.), or specify multiple points with the mouse button and press the rejection button. This part will be given a rejection label.
[0090] In one embodiment of the present invention, the method of designating incorrect parts in a surgical video and attaching a label is as follows.
[0091] First, limit the video segment in which an incorrect medical act and / or incorrect medical movement was performed in chronological order (or on the timeline within the corresponding playback bar) using the mouse cursor. The video information existing between the times selected by moving the corresponding mouse cursor is limited to the information for labeling.
[0092] In one embodiment of the present invention, when attempting to perform selection labeling on a video of implanting a corrective mini - implant, use a plurality of tags using mouse dragging and mouse buttons on the still image and / or video screen of the video to specify a point cloud that becomes a boundary line (including, for example, curves, straight lines, etc.) and / or a boundary surface (including, for example, closed curves, etc.). At this time, the video to be selected is automatically recognized, and an approval button can be pressed for the video recognized in the next order.
[0093] Further, the terminal 100 transmits to the server 200 one or more selection label values at one or more specific time points (or specific intervals) related to the load data, meta - information of the corresponding load data, identification information of the corresponding terminal 100, and the like.
[0094] Also, the terminal 100, in conjunction with the server 200, performs hierarchical labeling on the corresponding one or more load data before or after performing selection labeling on the corresponding one or more load data, and can also perform selection labeling on the corresponding one or more load data before / after performing hierarchical labeling. Here, the hierarchical labeling (or hierarchical leveling / primary hierarchical labeling / first - level hierarchical labeling) is a labeling method in which, by user input feature engineering (or hierarchical clustering labeling), a label (or label value) indicating a feature related to the corresponding load data is attached, and the corresponding load data is divided (or classified) into a plurality of sub - load data according to the feature.
[0095] That is, in conjunction with the server 200, the terminal 100 refers to (or is based on) a plurality of preset label classifications related to the corresponding specific topic for the load data displayed on the corresponding terminal 100, and according to the input of the user of the corresponding terminal 100 (or the user's selection / touch / control), it sets (or receives / inputs) the label (or label value) at another specific time point (or another specific section) of the corresponding load data.
[0096] The following [Table 1] to [Table 11] show examples of label classifications (or label classification tables) in a specific field.
[0097] The label classification shows a correct data set for artificial intelligence to learn, and shows a classification hierarchically in any manner and step so that the user can refer to it for hierarchical clustering labeling.
[0098] That is, the above [Table 1] to the above [Table 6] show examples of hierarchical label values (or variable values) in the process of a dental university professor (or doctor influencer) performing an implant operation or a laminating operation. At this time, for various operations, there are m1×m2×m3×n×n′×N (the product of each label classification) label classifications for the doctor influencer.
[0099] The user refers to the label classification and inputs variable values (or label values) related to hierarchical clustering into the input windows s1, s2, s3 of FIGS. 23 to 28 and FIGS. 30 to 32.
[0100] The variable values (or label values) of the variable s1 of the first hierarchies 201, 701, 801, 901, 1001, 1101 are input, the variable values (or label values) of the variable s2 of the second hierarchies 102, 702, 802, 902, 1002, 1102 are input, and the variable values (or label values) of the variable s3 of the third hierarchies 203, 703, 803, 903, 1003, 1103 are input. The number of input windows increases according to the number of hierarchies.
[0101] When a user referring to label classification moves a table (or cursor) indicating the time on the timeline within the playback bar (Figs. 23 to 28, Figs. 29 to 32) and captures the still image information at the point in time when trying to split the video, and then presses or selects the ACCEPT button, the corresponding selection time point becomes the video splitting time point, and the terminal 100 (or the server 200) splits the corresponding video according to the selected time point (or splitting time point) corresponding to the ACCEPT button.
[0102] Video information such as the fourth - level 204, 704, 804, 904, 1004, 1104, the fifth - level 905, 1005, 1105, and the sixth - level 1106 is split in the order of label values k, L, f of label classification.
[0103]
Table 1
[0104]
Table 2
[0105]
Table 3
[0106]
Table 4
[0107]
Table 5
[0108]
Table 6
[0109] Further, the said [Table 7] to the said [Table 11] show examples of hierarchical label values (or variable values) in the dance movements of As If It's Your Last by Blackpink for dancers (or dance influencers).
[0110]
Table 7
[0111]
Table 8
[0112]
Table 9
[0113]
Table 10
[0114]
Table 11
[0115] Thus, the said [Table 5] and the said [Table 10] are label classifications that are subcategorized into characteristic movements so that the video can be short - divided by the user around 1 to 3 seconds. The actual robot operation videos may be produced for label classification in the same manner as the said [Table 1] to [Table 12].
[0116] Also, the said [Table 6] and the said [Table 11] are those with labels attached to body parts of avatars, humans, robots, etc., and are used for first - and second - layer labeling, selective labeling, additional selective labeling, time - series segmentation selective labeling, selective labeling by body part, etc., and are label classifications arbitrarily set by an expert group.
[0117] Selection by body part is a method of labeling each video of a detailed body part with a label classification [Table 6] for each detailed body part of the body and assigning label values in the label order of the said [Table 11] through object recognition for each detailed body part in a single or multiple divided still images. The video can be divided in data unit 5 using the playback bar (table indicating time) or selection by body part.
[0118] Hierarchical labeling through selection by body part (generating data unit 5, video segmentation, f) is input feature engineering executed by the user, and such selection by body part may be omitted. The server 200 can call a library regarding selection by body part (such as object recognition of detailed body parts) to automatically label (label f) and divide the video (data unit 5).
[0119] FIG. 11 may be used as a diagram showing hierarchical clustering based on data unit 5.
[0120] Selection by body part is hierarchical labeling, and selection-by-body-part labeling is labeling that generates digital unit 5 and divides the video according to the interaction between the server 200 and the user (the user's judgment regarding the video segmentation time point (label value) by the server or the judgment regarding the operation order of body parts).
[0121] Time-series segmentation selection (generating data units 3 and 4) is hierarchical labeling, and time-series segmentation selection labeling is labeling that generates digital units 3 and 4 and divides the video according to the interaction of the user (the user's judgment regarding the video segmentation time point (label value) by the server).
[0122] The detailed operation steps are classified into detailed operation step 1 and detailed operation step 2 according to the method of dividing the video. Here, the said detailed operation step 1 is the detailed division of the operation steps by time-series segmentation selection labeling, and the said detailed operation step 2 is the detailed division of the operation steps by selection-by-body-part labeling.
[0123] In one embodiment of the present invention, the dance movements of the open concert of Jennie's song As If It's Your Last (3 minutes and 14 seconds), which is a Jennie song in the label classification [Table 9], broadcast on July 8, 2022, are output as images on the user's HMD (Head-mounted display). Jennie's image is in video form and is viewed by the user in a segmented form. The user can view the still images in label order and can view the still image at the end of the segmented video.
[0124] The user refers to the image of Jennie output by the HMD and executes movements similar to or identical to Jennie's movements on a VR treadmill. The movement image information of the user is collected by the visual set device and used as raw data (or basic video / basic video information). The user can also generate their own avatar synthesized with Jennie's movements according to their own choice, or can output their own appearance and movements as they are without being synthesized with Jennie's movements.
[0125] At this time, the user refers to the still images and label values of the still images that are seen as being in a stopped state in Jennie's movements, and performs hierarchical labeling, selective labeling, time-series segmentation selective labeling, selective labeling by body part, etc. on their own avatar and other people's avatars. The user can also view the movements of their own avatar and other people's avatars generated by being synthesized with Jennie's avatar by artificial intelligence on the HMD at a third-party time point, and perform the above labeling by comparing it with the dance movements of the open concert of Jennie's song As If It's Your Last (3 minutes and 14 seconds) broadcast on July 8, 2022 in the label classification [Table 9].
[0126] The user can imitate Jeni's movements several times, and the basic video information (or raw data), which is such dance movement, can be collected by the terminal 100 and transmitted to the server 200. The dance movement information for several times is a plurality of data (or a plurality of raw data) of an attribute item (or annotation step).
[0127] In one embodiment of the present invention, the above [Table 1] to [Table 6] show examples of hierarchical label values (or variable values) in the process of a dental university professor (or doctor influencer) performing an implant surgery or a laminating procedure. Dental university students and dentists can use a tooth removal VR simulator (not shown) to proceed with virtual surgeries, virtual procedures, etc. on a digital CAD while viewing the label classification of the above [Table 1] to [Table 6], which is a correct answer data set, by means of an HMD, and can proceed with labeling.
[0128] At this time, when the playback bar included in the view screen within the execution result screen of the app displayed on the terminal 100 is selected, or when the playback button within the corresponding view screen is selected, the terminal 100 displays (or outputs) the collected raw data in the video display area, and displays (or outputs) a comparison target video corresponding to the collected raw data (or a comparison target video corresponding to the corresponding raw data provided by the server 200) in the display area of the comparison target video. At this time, the terminal 100 may synchronize the corresponding raw data and the comparison target video based on the meta information corresponding to the raw data and the comparison target video respectively, and display the synchronized raw data and comparison target video in the video display area and the display area of the comparison target video respectively. Here, when either one of the raw data displayed in the video display area within the terminal 100 and the comparison target video displayed in the display area of the comparison target video stops due to a pause function or a stop function, the terminal 100 controls the other one to also stop by means of the pause function or the stop function.
[0129] In addition, for the load data displayed in the video display area of the terminal 100, the terminal 100 sets (or receives / inputs) one or more step-by-step labels (or label values) for the movement (or actions of the object) of the object included in the corresponding load data at other specific time points (or other specific intervals) according to the input of the user of the corresponding terminal 100 (or the user's selection / touch / control).
[0130] That is, at one or more other specific time points (or other specific intervals) of the load data displayed in the video display area, the terminal 100 hierarchically inputs hierarchical labels (or hierarchical label values) for the movement (or actions of the object) of the object included in the corresponding load data according to the user's input, with respect to specific operations of the object, specific ways of specific operations, specific steps of specific ways, etc.
[0131] In this way, for the load data related to the corresponding specific topic, the terminal 100 sets (or receives / inputs) one or more hierarchical labels (or hierarchical label values) at one or more other specific time points (or other specific intervals) according to the input of the user of the corresponding terminal 100, who is an expert related to the corresponding specific topic.
[0132] In addition, before / after performing the hierarchical labeling process, the terminal 100 performs the selection labeling process described above.
[0133] In this way, the terminal 100 performs hierarchical labeling, which is hierarchical clustering labeling for inputting label values with reference to the label classification that hierarchically classifies the actions of an avatar, a human, a robot, etc. according to the action type and / or specific way and / or specific step and / or step of detailed action of the specific avatar, human, robot, etc.
[0134] Actions related to the treatment or surgery of a patient with an anatomical structure (or specific mutation) similar to a certain specific patient are included in the specific actions of a specific avatar, human, robot, etc.
[0135] In the hierarchical clustering of dental procedures or surgical operations, the label classification of the specific methods for the specific cases in the above [Table 2] to [Table 3] is included in the label classification by specific actions and specific action methods of specific avatars, humans, robots, etc. in the above [Table 7] to [Table 9].
[0136] In one embodiment of the present invention, the artificial intelligence that has learned the label values of hierarchical labeling in the server 200 returns the label values or video information to the user of the terminal 100, and the user can attach an approval label or a rejection label thereto.
[0137] In various embodiments, the user may be an expert having an occupation in each field (including, for example, domain experts, dentists, doctors, soccer players, dancers, etc.).
[0138] The cuboid 301 in FIG. 3 is the video information of the segmented video which is the target attribute in FIGS. 4 to 6, and shows the video information 405 of the segmented action at the k-th step, the video information 505 of the segmented action at the L-th step, and the video information 605 of the segmented action at the f-th step. Here, the starting part yz plane 302 or the ending part yz plane 303 shows the still image information which is an attribute.
[0139] Here, the variables m1, m2, and m3 represent arbitrary positive integers (or natural numbers), s1, s2, and s3 represent variables, JPEG2025517145000013.jpg934, JPEG2025517145000014.jpg934, JPEG2025517145000015.jpg934 are shown, k represents a variable, JPEG2025517145000016.jpg926 is shown, the variables n, n′, and N represent arbitrary positive integers (or natural numbers), L and f represent variables, JPEG2025517145000017.jpg927 and JPEG2025517145000018.jpg926 are shown.
[0140] The user first checks on the screen displayed on the corresponding terminal 100 the output of the load data (or video information) corresponding to the first attribute and the first target attribute, and then inputs the hierarchical clustering related variable values (or label values) on the screen displayed on the corresponding terminal 100 with reference to the label classification.
[0141] In addition, the terminal 100 outputs video information (for example, including attributes, target attributes, etc.) related to the operations of avatars, humans, robots, etc.
[0142] In addition, the attributes and target attributes in FIGS. 4 to 6 may be virtual avatar, item, human, robot, etc. operation-related video information generated by the server 200.
[0143] In addition, the terminal 100 receives the variable value (label value) of the variable S1 of the first layer corresponding to the user's input through a plurality of input windows included in the hierarchical label input menu within the corresponding view screen, receives the variable value (or label value) of the variable S2 of the second layer, and receives the variable value (or label value) of the variable S3 of the third layer. At this time, the plurality of input windows included in the hierarchical label input menu can be variously set according to the number of layers according to the designer's design.
[0144] In addition, the video information of the fourth layer, the fifth layer, and the sixth layer is divided in the order of label values (for example, k, L, f) of the label classification.
[0145] In an embodiment of the present invention, the method for the user to split steps is to move the table (or cursor) indicating the time on the timeline in the playback bar with the mouse from the time axis, check the time of the video to be split, and then select it with the mouse.
[0146] In the embodiments of the present invention, although mainly described is that the user sets hierarchical labeling related to specific load data with reference to the label classification, the present invention is not limited thereto. The terminal 100 may also directly input hierarchical labeling values step by step (or hierarchically / in a cascade form) according to the input of the user of the corresponding terminal 100.
[0147] FIG. 4 shows an example of a state in which a video (or a long three-dimensional figure) is divided into n pieces, FIG. 5 shows an example of a state in which a video is divided into n' pieces, and FIG. 6 shows an example of a state in which a video is divided into N pieces. Here, n, n' and N satisfy JPEG2025517145000019.jpg926, JPEG2025517145000020.jpg927 and JPEG2025517145000021.jpg926, and the variables k, L and f represent positive integers (or natural numbers).
[0148] FIGS. 4 to 6 represent one operation of an avatar (or a human) in a three-dimensional figure. The first rectangle (shown in black) is the still image information 401, 501, 601 at the start part of the operation (or video), the last rectangle (shown in black) is the still image information 404, 504, 604 at the end part of the operation (or video), the x-axis is time, the yz plane (rectangle) is still image information, and the divided rectangular parallelepipeds indicate the divided videos.
[0149] One of the divided three-dimensional figures in FIGS. 4 to 6 corresponds to (or matches) the rectangular parallelepiped 301 in FIG. 3.
[0150] Also, one of the divided three-dimensional figures in FIG. 5 is a long rectangular parallelepiped divided into n' pieces, and one of the divided three-dimensional figures in FIG. 6 is a long rectangular parallelepiped divided into N pieces. The long bar-shaped rectangular parallelepipeds in FIGS. 4 to 6 show the video information related to the overall operation of the avatar in a three-dimensional figure.
[0151] Also, referring to FIG. 3, the end portion yz plane 303, which is a black rectangular form, is an attribute (or still image), and the rectangular parallelepiped 301 is a video of a divided operation (or target attribute).
[0152] Next, they are matched in the order of the parentheses. The still image information 402, 502, 602 of the (k, L, f)-th start portion in FIGS. 4 to 6 indicates the still image information 302 of the start portion in FIG. 3, and the still image information 403, 503, 603 of the (k, L, f)-th end portion indicates the still image information 303 of the end portion in FIG. 3. The still image information 402, 502, 602 of the (k, L, f)-th start portion and the still image information 403, 503, 603 of the (k - 1, L - 1, f - 1)-th end portion are the same. The still image information 403, 503, 603 of the (k, L, f)-th end portion in FIGS. 4 to 6 is an attribute, and the (k, L, f)-th steps 405, 505, 605 of the divided operation video are target attributes. In FIGS. 3 to 6, data unit 1 indicates the sum of the still image information 401, 501, 601 of the start portion of the entire video of one operation such as an avatar, a human, or a robot, and the still image information 404, 504, 604 of the end portion of the entire video. Data unit 2 indicates the sum from the still image information of the first step to the still image information of the last step of the divided operation video of an avatar, a human, a robot, etc.
[0153] Data unit 3 in FIG. 4 indicates the sum of the video information of the k-th step of the divided operation video such as an avatar, a human, or a robot operation, and the still image information of the end portion of the k-th step.
[0154] Data unit 4 in FIG. 5 indicates the sum of the video information of the L-th step of the divided operation video such as an avatar, a human, or a robot operation, and the still image information of the end portion of the L-th step.
[0155] Data unit 4 in FIG. 6 indicates the sum of the video information of the f-th step of the divided operation video such as an avatar, a human, or a robot operation, and the still image information of the end portion of the f-th step.
[0156] Figure 7 shows hierarchical clustering based on data units 1, 2, and 3, and Figure 9 shows hierarchical clustering based on data units 1, 2, 3, and 4.
[0157] The attribute in the data unit 3 is the still image information 403 at the end of the k-th step of the segmented motion video such as avatar, human, and robot motion, and shows the black rectangular shape in FIG. 4.
[0158] Also, the attribute in the data unit 4 is the still image information 503 at the end of the L-th step of the segmented motion video such as avatar, human, and robot motion, and shows the black rectangular shape in FIG. 5.
[0159] Also, the attribute in the data unit 5 is the still image information 603 at the end of the f-th step of the segmented motion video such as avatar, human, and robot motion, and shows the black rectangular shape in FIG. 6.
[0160] The data units 3, 4, 5, etc. used in the classification model and prediction model (or induction and / or inference model) according to the embodiment of the present invention may be the divided rectangular parallelepipeds 301 in FIG. 3 (the same applies to digital units).
[0161] FIG. 2 is a phylogenetic diagram 900 of hierarchical clustering for actual real data (or raw data), showing K1 clusters. The raw data (actual real data) includes robot motion video information collected by a visual set device. When using the robot motion video information as raw data, FIG. 22 may be used for robot training.
[0162] FIG. 7 or FIG. 9 is a phylogenetic diagram of hierarchical clustering created by label values assigned when video steps are segmented by data units.
[0163] FIG. 7 shows K2 clusters based on data unit 3, and FIG. 9 shows K4 clusters based on data unit 4.
[0164] In one embodiment of the present invention, the still image information at the start portion is also an attribute, which forms a data unit by adding with the target attribute, and is used for forward video generation and output in the direction of the algorithm.
[0165] In one embodiment of the present invention, the labeling method related to hierarchical clustering by case, method, and step is as follows.
[0166] The above [Table 1] to [Table 6] are produced based on the specialized knowledge in the medical fields of doctors and dentists, and are examples of label classifications presented for inputting variable values (or label values) on the execution result screen (or view screen) of the application displayed on the terminal 100.
[0167] The above [Table 1] is an example for inputting variable values (or label values) in the surgical field, the above [Table 2] is an example for inputting variable values (or label values) in surgical cases, the above [Table 3] is an example for inputting variable values (or label values) in surgical methods, and the above [Table 4] is an example for inputting variable values (or label values) in surgical steps. This is a method for the user (including, for example, doctors and dentists) to refer to clinical criteria (including, for example, cases, methods, steps, etc.) and input variable values (or label values).
[0168] The above [Table 5] is an example of further subdividing and classifying surgical steps, and the above [Table 6] is an example of classification in which label values are specified for body parts such as avatars, humans, and robots.
[0169] Classification criteria for cases, methods, and steps for surgical operations are applied based on the data units in FIGS. 7 and 9 above by inputting hierarchical cluster label values. Although it is based on performing hierarchical clustering of medical video information and other information by case, method, and step, if the above information is detailedly labeled in an arbitrary manner, the video is segmented, and hierarchical clustering is performed, it can be applied to a classification model and / or a prediction model regardless of the type and number of hierarchies (including, for example, 3 layers, 4 layers, 5 layers, etc.) and the method of video segmentation.
[0170] In one embodiment of the present invention, video information of a patient's body and organs, other medical information, digital devices, etc. used in actual surgeries and procedures are labeled by the execution result screen (or view screen) of the application displayed on the terminal 100, and in FIGS. 7 to 10 above, they become K2, K3, K4, and K5 clusters respectively. Here, the K indicates a variable (or natural number). The body and organ information, digital devices, etc. of a specific patient belonging to the same cluster are the metadata (or meta-information) corresponding to that cluster. Virtual surgeries, virtual procedures, etc. are advanced by artificial intelligence inference or restoration of digital devices based on the metadata. For virtual surgery videos, virtual procedure videos, etc. utilizing the digital device and the artificial device output by artificial intelligence inference and restoration, doctors and / or dentists perform selective labeling through the execution result screen (or view screen) of the application displayed on the terminal 100.
[0171] In various embodiments, in FIGS. 3 to 6 above, one operation of the avatar (or human) is one surgical operation of a specific patient. The first still video information at the start of the operation is diagnostic information, and the steps of the k, L, and f-th operations of the video information of one operation of the avatar (or human) are the steps of the k, L, and f-th surgeries of one surgical operation video of a specific patient. When comparing the reaction of the digital device undergoing the surgery with the operation of the avatar or human, it can be said to be a kind of passive operation of the avatar (for example, the digital device is a kind of patient avatar).
[0172] In this way, the terminal 100 performs functions such as hierarchical labeling and selective labeling on images of avatars, items, robots, etc.
[0173] In the embodiments of the present invention, the hierarchical labeling function and the selective labeling function are described separately, but the present invention is not limited thereto. The terminal 100 can perform the hierarchical labeling function by including it in the selective labeling function, and can also integrate the hierarchical labeling and the selective labeling into one labeling function.
[0174] Also, the terminal 100 receives a first video transmitted from the server 200. Here, the first video is a result generated by the learning results of a classification model and a prediction model for the corresponding load data in the server 200, and may be an operation-related video of an avatar, an item, a robot, etc. generated based on the load data, a video in which the load data is updated (for example, a video in which the actions / behavior / conduct of a human / humans included in the load data are updated), etc.
[0175] Also, the terminal 100 outputs the received first video to the video display area. At this time, the terminal 100 can also divide the screen of the corresponding terminal 100 and output it simultaneously in a state where the load data, the comparison target video, and the first video are synchronized.
[0176] Also, the terminal 100 performs additional selection labeling on the basis of the first video in conjunction with the server 200. Here, the additional selection labeling (or additional selection labeling / secondary selection labeling / second selection labeling) refers to a labeling method for setting (or attaching) a label (or label value) for the presence or absence of an error (or abnormality) at another specific time point (or another specific section) of the first video. At this time, for a time point (or section) of the first video for which a label (or label value) has not been set by the additional selection labeling, a preset default label value (for example, an approval label) may be set. Further, the terminal 100 attaches a preset not ACCEPT label to a time point (or section / attribute / target attribute) of the first video for which the approval label is not attached, and may also attach a preset not REJECT label to a time point (or section / attribute / target attribute) of the first video for which the rejection label is not attached.
[0177] That is, the terminal 100, in conjunction with the server 200, sets (or receives / inputs) a label (or label value) for the presence or absence of an error (or abnormality) at another specific time point (or another specific section) of the first video displayed on the corresponding terminal 100 in response to an input (or user selection / touch / control) of the user of the corresponding terminal 100.
[0178] Also, when the playback bar included in the view screen within the execution result screen of the app displayed on the terminal 100 is selected, or when the playback button within the corresponding view screen is selected, the terminal 100 displays (or outputs) the first video in the video display area, and displays (or outputs) a comparison target video corresponding to the load data (or the first video) (or a comparison target video corresponding to the corresponding load data / first video provided by the server 200) in the display area of the comparison target video. At this time, the terminal 100 may perform synchronization on the corresponding first video and the comparison target video based on the meta information corresponding to the first video and the comparison target video respectively, and display the synchronized first video and comparison target video in the video display area and the display area of the comparison target video respectively. Here, when either one of the first video displayed in the video display area within the terminal 100 and the comparison target video displayed in the display area of the comparison target video stops due to a pause function or a stop function, the terminal 100 controls the other one to also stop by the pause function or the stop function.
[0179] Also, the terminal 100 sets (or receives / inputs) a label (or label value) for a correct action or an incorrect action with respect to the movement (or the action of the object / avatar) of the object (or avatar) included in the corresponding first video at another specific time point (or another specific section) in response to the input (or the selection / touch / control of the user) of the user of the terminal 100 with respect to the first video displayed in the video display area of the terminal 100.
[0180] That is, at one or more other specific time points of the first video displayed in the video display area, the terminal 100 inputs a label value for a correct action (for example, a preset approval / acceptance / ACCEPT label) or a label value for an incorrect action (for example, a preset rejection / REJECT label) in response to the input of the user.
[0181] In this way, for the first video generated in relation to the corresponding specific topic, the terminal 100 sets (or receives / inputs) one or more additional selection labels (or additional selection label values) at one or more other specific time points (or other specific intervals) according to the input of the user of the terminal 100 who is an expert related to the corresponding specific topic.
[0182] At this time, the terminal 100 performs a time-series segmentation selection labeling function or a selection labeling function by body part according to the input of the user of the terminal 100.
[0183] The terminal 100 performs the time-series segmentation selection labeling function through the following process.
[0184] That is, for the plurality of sub-videos obtained by splitting the first video, according to the input of the user, label values for the correct state (or correct action) of the split state of each sub-video (for example, a preset approval / acceptance / ACCEPT label) or label values for an incorrect state (or incorrect action) (for example, a preset rejection / REJECT label) are respectively input. In order to sort the order of the corresponding plurality of sub-videos, according to the input of the user, a label value indicating the order of the corresponding plurality of sub-videos (or a label value for adjusting the split time point if the split time point is incorrect or adjustment is required) is input. Here, the splitting of the first video into a plurality of sub-videos may be in a state where the first video is split into the plurality of sub-videos based on information regarding the plurality of sub-data split by performing the hierarchical labeling function on the raw data, or may be in a state where the first video is split into the plurality of sub-videos by performing an artificial intelligence function or a video analysis function on the raw data in the server 200.
[0185] As a result, for the first video, according to the input of the user of the corresponding terminal 100, label values for the correct state and the incorrect state of the division state of the plurality of sub-videos are respectively input, and label values for sorting the order of the corresponding plurality of sub-videos (or label values indicating the order of the corresponding plurality of sub-videos / if the division time point is incorrect or adjustment is required, label values for adjusting the division time point) are respectively input.
[0186] In addition, the terminal 100 transmits to the server 200 the label values for the correct state and the incorrect state of the division state of the input plurality of sub-videos, the label values for sorting the order of the plurality of sub-videos (or label values for adjusting the division time point), the identification information of the terminal 100, and the like.
[0187] In addition, the server 200 receives, by performing the time-series division selection labeling function for the corresponding first video, the label values for the correct state and the incorrect state of the division state of the plurality of sub-videos transmitted from the terminal 100, the label values for sorting the order of the plurality of sub-videos (or label values for adjusting the division time point), the identification information of the terminal 100, and the like.
[0188] In addition, the server 200 reorders the order of the plurality of sub-videos obtained by dividing the first video based on the label values for the correct state and the incorrect state of the division state of the received plurality of sub-videos, the label values for sorting the order of the plurality of sub-videos (or label values for adjusting the division time point), and the like.
[0189] In this way, when the first video is divided into a plurality of sub-videos, the time-series division selection labeling may be a process of labeling whether each division time point (for example, including label values, still image information, etc.) of the plurality of sub-videos of the corresponding first video is correct or incorrect according to the input of the user of the terminal 100, and labeling the label values for adjusting the division time point or the order when the division time point is incorrect.
[0190] Further, the terminal 100 performs a selection labeling function for each body part through the following process.
[0191] That is, for the avatars (or objects) included in the plurality of sub-images obtained by dividing the first video, the terminal inputs label values for the operation order of the avatars (or objects) included in the plurality of sub-images (or label values for the correct or incorrect state of the operation order of the corresponding avatar) according to the user's input. In order to align the operation order for each body part (or each part of the robot) from the operations of the avatars, humans, robots, etc. included in the corresponding plurality of sub-images, according to the user's input, a label value indicating the order of the corresponding plurality of sub-images (or a label value for adjusting the order of the sub-images including the avatar) is input. Such selection for each body part may be executed or omitted by the user, or may be automatically executed by the server 200 (hierarchical labeling). Here, the division of the first video into a plurality of sub-videos is based on information regarding the sub-data divided into a plurality by performing the hierarchical labeling function for the raw data, and the first video is in a state of being divided into the plurality of sub-videos, or the first video may be in a state of being divided into the plurality of sub-videos by performing the artificial intelligence function or video analysis function for the raw data in the server 200.
[0192] Thereby, for the first video, according to the user input of the corresponding terminal 100, label values for the operation order of the avatars (or objects) included in the plurality of sub-images (or label values for the correct or incorrect state of the operation order of the corresponding avatar, robot, etc.) are respectively input, and label values for aligning the order of the corresponding plurality of sub-images (or the operation order of the avatars, robots, etc. included in the corresponding plurality of sub-images) (or label values indicating the order of the corresponding plurality of sub-images / label values for adjusting the order of the sub-images including the avatar, robot) are respectively input.
[0193] Further, the terminal 100 transmits to the server 200 a label value for the operation order of the avatar (or object) included in the plurality of input sub-videos (or a label value for the correct or incorrect state of the operation order of the corresponding avatar, robot, etc.), a label value for aligning the order of the plurality of sub-videos (or the operation order of the avatar, robot, etc. included in the corresponding plurality of sub-videos), identification information of the terminal 100, and the like.
[0194] Further, the server 200 receives, by performing a selection labeling function for each body part for the corresponding first video, a label value for the operation order of the avatar (or object) included in the plurality of sub-videos transmitted from the terminal 100 (or a label value for the correct or incorrect state of the operation order of the corresponding avatar), a label value for aligning the order of the plurality of sub-videos (or the operation order of the avatar included in the corresponding plurality of sub-videos), identification information of the terminal 100, and the like.
[0195] Further, the server 200 re-aligns the order of the plurality of sub-videos obtained by splitting the first video based on a label value for the operation order of the avatar (or object) included in the received plurality of sub-videos (or a label value for the correct or incorrect state of the operation order of the corresponding avatar), a label value for aligning the order of the plurality of sub-videos (or the operation order of the avatar included in the corresponding plurality of sub-videos), and the like.
[0196] In this way, when the first video is split into a plurality of sub-videos, the selection labeling for each body part may be a process of labeling whether the operation order of the avatar (or object) included in each of the plurality of split sub-videos of the corresponding first video is correct or incorrect in response to an input from the user of the terminal 100, and labeling a label value for aligning the order of the plurality of sub-videos (or the operation order of the avatar included in the corresponding plurality of sub-videos) in order to adjust the operation order of the corresponding avatar.
[0197] In addition, the selection labeling function for each body part further includes the following functions.
[0198] That is, the server 200 provides the terminal 100 with information on the operation sequence of avatars, robots, etc. included in the plurality of sub-images with respect to the plurality of divided sub-images by performing the artificial intelligence function and video analysis function in the server 200.
[0199] Also, in the terminal 100, in response to a user's input, the terminal 100 labels the operation sequence of the avatar (or robot) included in the corresponding plurality of sub-images as correct or incorrect. When the operation sequence of the avatar (or robot) is incorrect or needs adjustment, label values for adjusting the operation sequence or the order of the sub-images including the avatar or robot are input. The label value for the correct or incorrect state with respect to the operation sequence of the avatar (or robot) included in the corresponding plurality of input sub-images, the label value for adjusting the operation sequence or the order of the sub-images including the avatar (or robot) when the operation sequence of the avatar (or robot) is incorrect or needs adjustment, and the identification information of the terminal 100 are transmitted to the server 200.
[0200] In addition, the server 200 receives the label value for the correct or incorrect state with respect to the operation sequence of the avatar (or robot) included in the corresponding plurality of sub-images transmitted from the terminal 100, the label value for adjusting the operation sequence or the order of the sub-images including the avatar (or robot) when the operation sequence of the avatar (or robot) is incorrect or needs adjustment, and the identification information of the terminal 100.
[0201] In addition, the server 200 re-arranges the order of a plurality of sub-images obtained by splitting the first image based on label values for correct or incorrect states with respect to the operation order of the avatar (or robot) included in the received corresponding plurality of sub-images, label values for adjusting the order of the operation order or the sub-images including the avatar (or robot) when the operation order of the avatar (or robot) is incorrect or needs adjustment, and the like.
[0202] In addition, the terminal 100 transmits to the server 200 one or more additional selection label values, one or more time-series split selection label values, one or more selection label values for each body part, label values for arranging the order of a plurality of sub-images, identification information of the corresponding terminal 100, and the like at one or more other specific time points (or one or more other specific sections) related to the first image.
[0203] In addition, the terminal 100, in conjunction with the server 200, performs additional hierarchical labeling on one or more corresponding first images before or after performing additional selection labeling on the corresponding first image, and can also perform additional selection labeling on the corresponding first image before / after performing the additional hierarchical labeling. Here, the additional hierarchical labeling (or additional hierarchical leveling / secondary hierarchical labeling / second hierarchical labeling) is user input feature engineering, and indicates a labeling method of attaching a label (or label value) indicating a feature of the corresponding first image and splitting (or classifying) the corresponding first image into a plurality of sub-images according to the feature.
[0204] That is, the terminal 100, in conjunction with the server 200, refers to (or based on) a plurality of preset label classifications related to a specific topic for the first image displayed on the corresponding terminal 100, and sets (or receives / inputs) additional labels (or additional label values) at one or more other specific time points (or one or more other specific sections) of the corresponding first image according to the input (or user selection / touch / control) of the user of the corresponding terminal 100.
[0205] At this time, when the playback bar included in the view screen within the execution result screen of the application displayed on the terminal 100 is selected, or when the playback button within the corresponding view screen is selected, the terminal 100 displays (or outputs) the first video in the video display area, and displays (or outputs) a comparison target video related to the first video (or the corresponding comparison target video provided by the server 200 for the relevant load data / first video) in the display area of the comparison target video. At this time, the terminal 100 may perform synchronization on the corresponding first video and the comparison target video based on the meta information corresponding to the first video and the comparison target video respectively, and display the synchronized first video and comparison target video in the video display area and the display area of the comparison target video respectively. Here, if either one of the first video displayed in the video display area within the terminal 100 and the comparison target video displayed in the display area of the comparison target video stops due to a pause function or a stop function, the terminal 100 controls the other one to also stop by the pause function or the stop function.
[0206] In addition, the terminal 100 sets (or receives / inputs) one or more step-by-step additional labels (or additional label values) for the movement (or behavior of the object) of the object included in the first video displayed in the video display area of the terminal 100 in response to the input of the user of the corresponding terminal 100 (or the user's selection / touch / control), and also at other specific time points (or other specific intervals).
[0207] That is, at one or more other specific time points (or other specific intervals) of the first video displayed in the video display area, the terminal 100, in response to the input of the user, hierarchically inputs additional hierarchical labels (or additional hierarchical label values) for the movement (or behavior of the object) of the object included in the corresponding first video, such as specific operations of the object, specific methods of the specific operations, and specific steps of the specific methods.
[0208] In this way, for the first video related to the corresponding specific topic, the terminal 100 sets (or receives / inputs) one or more additional hierarchical labels (or additional hierarchical label values) at one or more other specific time points (or other specific intervals) in response to the input of the user of the terminal 100, who is an expert related to the corresponding specific topic.
[0209] Also, before / after performing the additional hierarchical labeling process, the terminal 100 performs the additional selection labeling process described above.
[0210] In this way, the terminal 100 performs functions such as an additional hierarchical labeling function and an additional selection labeling function on the first video.
[0211] In the embodiments of the present invention, the additional hierarchical labeling function and the additional selection labeling function are separately described, but the present invention is not limited thereto. The terminal 100 can perform the additional hierarchical labeling function by including it in the additional selection labeling function, and can also integrate the additional hierarchical labeling and the additional selection labeling into one additional labeling function.
[0212] Also, the terminal 100 receives a second video transmitted from the server 200. Here, the second video is a result generated by a classification model and a prediction model for the corresponding first video in the server 200, and may be an avatar, an item, an operation-related video of a robot, etc. generated based on the first video, or a video obtained by updating the first video.
[0213] Also, the terminal 100 outputs the received second video to the video display area. At this time, the terminal 100 can also divide the screen of the corresponding terminal 100 and output it simultaneously in a state where the raw data, the comparison target video, the first video, and the second video are synchronized.
[0214] Further, the terminal 100 may be provided with a second video that has been newly aggregated and intellectualized (or an updated second video) related to the specific topic (or the downloaded data) from the server 200.
[0215] In addition, the terminal 100 transmits, to the server 200, operation-related videos of an avatar, an item, a robot, etc. (or operation-related videos related to at least one of the avatar and the item) output from the terminal 100, meta information related to the corresponding operation-related videos, etc., in relation to a specific topic. Here, the specific topic (or specific content) includes medical acts (including, for example, treatments, surgeries, etc.), dance, sports events (including, for example, soccer, basketball, table tennis, etc.), games, e-sports, etc. Also, the operation-related videos of the avatar and / or the item may be videos generated by a selection labeling process, a classification model inference process, a prediction model inference process, etc. based on any downloaded data related to the corresponding specific topic. The robot video is a video (or downloaded data) obtained by collecting actual robot operations with a visual set device.
[0216] In addition, in conjunction with the server 200, the terminal 100 sets (or receives / inputs) a label (or label value) at a specific point in time (or specific section) of the corresponding robot operation video in response to an input (or user selection / touch / control) of a user of the corresponding terminal 100 for the robot operation video (Fig. 29, basic robotics video) displayed on the corresponding terminal 100.
[0217] Also, when the playback bar included in the view screen within the execution result screen of the app displayed on the terminal 100 is selected, or when the playback button within the corresponding view screen is selected, the terminal 100 displays (or outputs) the robot operation video in the video display area, and displays (or outputs) the comparison target video corresponding to the robot operation video (or the comparison target video corresponding to the corresponding robot operation video provided by the server 200) in the display area of the comparison target video. At this time, the terminal 100 synchronizes the corresponding robot operation video and the comparison target video based on the meta information corresponding to the robot operation video and the comparison target video respectively, and may display the synchronized robot operation video and comparison target video in the video display area and the display area of the comparison target video respectively. Here, when either one of the robot operation video displayed in the video display area within the terminal 100 and the comparison target video displayed in the display area of the comparison target video stops due to a pause function or a stop function, the terminal 100 controls the other one to also stop by the pause function or the stop function.
[0218] Also, the terminal 100 sets (or receives / inputs) a label (or label value) for the correct action or incorrect action with respect to the movement (or action of the object) of the object included in the corresponding robot operation video at a specific time point (or specific section) in response to the input (or user's selection / touch / control) of the user of the corresponding terminal 100 for the robot operation video displayed in the video display area of the terminal 100.
[0219] That is, at one or more specific time points of the robot operation video displayed in the video display area, the terminal 100 inputs a label value for the correct action (for example, a preset approval / acceptance / ACCEPT label) or a label value for the incorrect action (for example, a preset rejection / REJECT label) respectively in response to the user's input.
[0220] As described above, for the robot operation video related to the corresponding specific topic, the terminal 100 sets (or receives / inputs) one or more selection labels (or selection label values) at one or more specific time points (or specific intervals) in response to the input of the user of the terminal 100, who is an expert related to the corresponding specific topic. Here, the selection labeling (or selection re-labeling / primary selection labeling / first selection labeling) refers to a labeling method of setting (or attaching) a label (or label value) for the presence or absence of an error (or abnormality) at a specific time point (or specific interval) of the robot operation video. At this time, for the time point (or interval) of the robot operation video where no label (or label value) is set by the selection labeling, a preset default label value (for example, an approval label) may be set. In addition, the terminal 100 attaches a preset not ACCEPT label to the time point (or interval / attribute / target attribute) of the robot operation video that does not have the approval label, and can also attach a preset not REJECT label to the time point (or interval / attribute / target attribute) of the robot operation video that does not have the rejection label.
[0221] In addition, the terminal 100 transmits one or more selection label values at one or more specific time points (or specific intervals) related to the robot operation video, the meta information of the corresponding robot operation video, the identification information of the terminal 100, etc. to the server 200.
[0222] In addition, the terminal 100, in conjunction with the server 200, performs hierarchical labeling on the corresponding robot operation video before or after performing selection labeling on the corresponding robot operation video, and can also perform selection labeling on the corresponding robot operation video before / after performing hierarchical labeling. Here, the hierarchical labeling (or hierarchical re-labeling) is user-based input feature engineering, which refers to a labeling method of attaching a label indicating the feature of the corresponding robot operation video and dividing (or classifying) the corresponding robot operation video into a plurality of sub-robot operation videos according to the feature.
[0223] That is, the terminal 100, in conjunction with the server 200, refers to (or is based on) a plurality of preset label classifications related to the corresponding specific topic for the robot operation video displayed on the corresponding terminal 100, and according to the input of the user of the corresponding terminal 100 (or the user's selection / touch / control), sets (or receives / inputs) the label (or label value) at another specific time point (or another specific section) in the corresponding robot operation video.
[0224] At this time, when the playback bar included in the view screen within the execution result screen of the application displayed on the terminal 100 is selected, or when the playback button within the corresponding view screen is selected, the terminal 100 displays (or outputs) the robot operation video in the video display area, and displays (or outputs) the comparison target video corresponding to the robot operation video (or the comparison target video corresponding to the corresponding robot operation video provided by the server 200) in the display area of the comparison target video. At this time, the terminal 100 may perform synchronization on the corresponding robot operation video and the comparison target video based on the meta information corresponding to the robot operation video and the comparison target video respectively, and display the synchronized robot operation video and comparison target video in the video display area and the display area of the comparison target video respectively. Here, when either one of the robot operation video displayed in the video display area within the terminal 100 and the comparison target video displayed in the display area of the comparison target video stops due to a pause function or a stop function, the terminal 100 controls the other one to stop by means of the pause function or the stop function.
[0225] In addition, the terminal 100, according to the input of the user of the corresponding terminal 100 (or the user's selection / touch / control), sets (or receives / inputs) one or more step-by-step labels (or label values) for the movement (or action of the object) of the object included in the corresponding robot operation video at another specific time point (or another specific section) with respect to the robot operation video displayed in the video display area of the terminal 100.
[0226] That is, at one or more other specific time points (or other specific intervals) of the robot operation video displayed in the video display area, the terminal 100, in response to a user input, hierarchically inputs hierarchical labels (or hierarchical label values) for the movement (or actions of the object) of the object included in the corresponding robot operation video with respect to specific operations of the object, specific ways of specific operations, specific steps of specific ways, and the like.
[0227] In this way, for the robot operation video related to the corresponding specific topic, the terminal 100 sets (or receives / inputs) one or more hierarchical labels (or hierarchical label values) at one or more other specific time points (or other specific intervals) in response to the input of the user of the terminal 100 who is an expert related to the corresponding specific topic.
[0228] Also, the terminal 100 performs the selection labeling process described above before / after performing the hierarchical labeling process.
[0229] Further, the terminal 100 receives the first robotics video transmitted from the server 200. Here, the first robotics video is a result generated by a classification model and a prediction model for the corresponding robot operation video in the server 200, and may be an operation-related video of an avatar, an item, a robot, etc. generated based on the robot operation video, a video in which the raw data is updated (for example, a video in which the movement / action / behavior of a human / humans included in the raw data is updated), and the like.
[0230] Also, the terminal 100 outputs the received first robotics video to the video display area. At this time, the terminal 100 can also divide the screen of the corresponding terminal 100 and output them simultaneously in a synchronized state of the robot operation video, the comparison target video, and the first robotics video.
[0231] Also, the terminal 100, in conjunction with the server 200, performs additional selection labeling on the first robotics video. Here, the additional selection labeling (or additional selection labeling / secondary selection labeling / second selection labeling) refers to a labeling method for setting (or attaching) a label (or label value) for the presence or absence of an error (or abnormality) at another specific time point (or another specific section) of the first robotics video. At this time, for the time point (or section) of the first robotics video for which a label (or label value) has not been set by the additional selection labeling, a preset default label value (for example, an approval label) may be set. Also, the terminal 100 attaches a preset not ACCEPT label to the time point (or section / attribute / target attribute) of the first robotics video that does not have the approval label, and can also attach a preset not REJECT label to the time point (or section / attribute / target attribute) of the first robotics video that does not have the rejection label. The second selection labeling may correspond to the first robotics selection labeling in FIG. 19.
[0232] That is, the terminal 100, in conjunction with the server 200, according to the input (or user selection / touch / control) of the user of the corresponding terminal 100, sets (or receives / inputs) a label (or label value) at another specific time point (or another specific section) of the corresponding first robotics video for the first robotics video displayed on the corresponding terminal 100.
[0233] Also, when the playback bar included in the view screen within the execution result screen of the application displayed on the terminal 100 is selected, or when the playback button within the corresponding view screen is selected, the terminal 100 displays (or outputs) the first robotics video in the video display area, and displays (or outputs) a comparison target video corresponding to the robot operation video (or the first robotics video) (or the comparison target video corresponding to the corresponding robot operation video / first robotics video provided by the server 200) in the display area of the comparison target video. At this time, the terminal 100 may synchronize the corresponding first robotics video and the comparison target video based on the meta information corresponding to the first robotics video and the comparison target video respectively, and display the synchronized first robotics video and comparison target video in the video display area and the display area of the comparison target video respectively.
[0234] Also, the terminal 100 sets (or receives / inputs) a label (or label value) for the correct behavior or incorrect behavior of the movement (or the actions of the object / avatar) of the object (or avatar) included in the corresponding first robotics video at another specific time point (or another specific section) in response to the input (or user's selection / touch / control) of the user of the terminal 100 with respect to the first robotics video displayed in the video display area of the terminal 100.
[0235] That is, at one or more other specific time points of the first robotics video displayed in the video display area, the terminal 100 inputs a label value for the correct behavior (for example, a preset approval / acceptance / ACCEPT label) or a label value for the incorrect behavior (for example, a preset rejection / REJECT label) in response to the user's input.
[0236] As described above, for the first robotics video generated in relation to the corresponding specific topic, the terminal 100 sets (or receives / inputs) one or more additional selection labels (or additional selection label values) at one or more other specific time points (or other specific intervals) according to the input of the user of the terminal 100 who is an expert related to the corresponding specific topic.
[0237] At this time, the terminal 100 performs a time-series segmentation selection labeling function or a selection labeling function by body part according to the input of the user of the terminal 100.
[0238] The terminal 100 performs the time-series segmentation selection labeling function through the following process.
[0239] That is, for the plurality of sub-robotics videos obtained by splitting the first robotics video, according to the input of the user, a label value for a correct state (or correct action) of the splitting state of each sub-robotics video (for example, a preset approval / acceptance / ACCEPT label) or a label value for an incorrect state (or incorrect action) (for example, a preset rejection / REJECT label) is input respectively. In order to sort the order of the corresponding plurality of sub-robotics videos, according to the input of the user, a label value indicating the order of the corresponding plurality of sub-robotics videos (or a label value for adjusting the splitting time point if the splitting time point is incorrect or adjustment is required) is input respectively. Here, the splitting of the first robotics video into a plurality of sub-robotics videos may be in a state where the first robotics video is split into the plurality of sub-robotics videos based on information on the plurality of sub-robotics videos obtained by performing the hierarchical labeling function on the robot operation video, or may be in a state where the first robotics video is split into the plurality of sub-robotics videos by performing an artificial intelligence function or a video analysis function on the robot operation video in the server 200.
[0240] As a result, for the first robotics video, the terminal 100 receives label values for the correct and incorrect split states of the plurality of sub-robotics videos respectively according to the input of the user of the corresponding terminal 100, and receives label values (or label values indicating the order of the corresponding plurality of sub-robotics videos / label values for adjusting the split time point if the split time point is incorrect or needs adjustment) for sorting the order of the corresponding plurality of sub-robotics videos respectively.
[0241] In addition, the terminal 100 transmits to the server 200 the label values for the correct and incorrect split states of the input plurality of sub-robotics videos, the label values for sorting the order of the plurality of sub-robotics videos (or label values indicating the order of the corresponding plurality of sub-robotics videos / label values for adjusting the split time point), the identification information of the terminal 100, etc.
[0242] In addition, the server 200 receives, by performing the time-series split selection labeling function for the corresponding first robotics video, the label values for the correct and incorrect split states of the plurality of sub-robotics videos transmitted from the terminal 100, the label values for sorting the order of the plurality of sub-robotics videos (or label values indicating the order of the corresponding plurality of sub-robotics videos / label values for adjusting the split time point), the identification information of the terminal 100, etc.
[0243] In addition, the server 200 re-sorts the order of the plurality of sub-robotics videos obtained by splitting the first robotics video based on the label values for the correct and incorrect split states of the received plurality of sub-robotics videos, the label values for sorting the order of the plurality of sub-robotics videos (or label values indicating the order of the corresponding plurality of sub-robotics videos / label values for adjusting the split time point), etc.
[0244] In this way, when the time-series segmentation selection labeling is performed and the first robotics video is divided into a plurality of sub-robotics videos, according to the input of the user of the terminal 100, it labels whether each segmentation time point (for example, including label values, still image information, etc.) of the plurality of sub-robotics videos corresponding to the relevant first robotics video is correct or incorrect. Even in the process of labeling the label values for adjusting the segmentation time point or order when the segmentation time point is incorrect, it may be such a process.
[0245] In addition, the terminal 100 performs a selection labeling function for each body part through the following process.
[0246] That is, for the avatars (or objects) included in the plurality of robotics sub-videos obtained by dividing the first robotics video, according to the input of the user, label values (or label values for the correct or incorrect state of the operation order of the corresponding avatar) for the operation order of the avatars (or objects) included in the plurality of sub-robotics videos are respectively input. In order to align the operation order for each body part from the operations of the avatars (or objects) included in the corresponding plurality of sub-robotics videos, according to the input of the user, label values indicating the order of the corresponding plurality of sub-robotics videos (or label values for adjusting the order of the sub-robotics videos including the avatar) are input. Here, the division of the first robotics video into a plurality of sub-robotics videos may be in a state where the first robotics video is divided into the plurality of sub-robotics videos based on information regarding the plurality of sub-robotics data obtained by performing the hierarchical labeling function for the robot operation video, or may be in a state where the first robotics video is divided into the plurality of sub-robotics videos by performing the artificial intelligence function or video analysis function for the robot operation video in the server 200.
[0247] As a result, for the first robotics video, the terminal 100 receives label values for the operation order of the avatars (or objects) included in the plurality of sub-robotics videos in response to the input of the user of the corresponding terminal 100 (or label values for the correct or incorrect state of the operation order of the corresponding avatar), and label values for aligning the order of the corresponding plurality of sub-robotics videos (or the operation order of the avatars included in the corresponding plurality of sub-robotics videos) (or label values indicating the order of the corresponding plurality of sub-robotics videos / label values for adjusting the order of the sub-robotics videos including the avatar) are respectively input.
[0248] In addition, the terminal 100 transmits to the server 200 the label values for the operation order of the avatars (or objects) included in the plurality of input sub-robotics videos (or label values for the correct or incorrect state of the operation order of the corresponding avatar), the label values for aligning the order of the plurality of sub-robotics videos (or the operation order of the avatars included in the corresponding plurality of sub-robotics videos), the identification information of the terminal 100, and the like.
[0249] In addition, the server 200 receives, by performing the body part-specific selection labeling function for the corresponding first robotics video, the label values for the operation order of the avatars (or objects) included in the plurality of sub-robotics videos transmitted from the terminal 100 (or label values for the correct or incorrect state of the operation order of the corresponding avatar), the label values for aligning the order of the plurality of sub-robotics videos (or the operation order of the avatars included in the corresponding plurality of sub-robotics videos), the identification information of the terminal 100, and the like.
[0250] Further, the server 200 re-aligns the order of the plurality of sub-robotics videos obtained by splitting the first robotics video based on a label value for the operation order of the avatar (or object) included in the received plurality of sub-robotics videos (or a label value for the correct or incorrect state of the operation order of the corresponding avatar), a label value for aligning the order of the plurality of sub-robotics videos (or the operation order of the avatar included in the corresponding plurality of sub-robotics videos), or a label value indicating the order of the corresponding plurality of sub-robotics videos.
[0251] In this way, when the first robotics video is split into a plurality of sub-robotics videos, the body part selection labeling labels whether the operation order of the avatar (or object) included in each of the plurality of sub-robotics videos obtained by splitting the corresponding first robotics video is correct or incorrect according to the input of the user of the terminal 100, and in order to adjust the operation order of the corresponding avatar, labels the label value for aligning the order of the plurality of sub-robotics videos (or the operation order of the avatar included in the corresponding plurality of sub-robotics videos).
[0252] In addition, the body part selection labeling function further includes the following functions.
[0253] That is, the server 200 provides the terminal 100 with information on the operation order of the avatar included in the plurality of sub-robotics videos by performing an artificial intelligence function or a video analysis function in the server 200 on the plurality of split sub-robotics videos.
[0254] Also, the terminal 100 labels, according to the user input, the operation order of the avatar included in the corresponding plurality of sub-robotics videos as being in a correct state or an incorrect state. When the operation order of the avatar is incorrect or needs adjustment, label values for adjusting the operation order or the order of the sub-robotics videos including the avatar (or human) are input. The terminal 100 transmits to the server 200 label values (e.g., including selection, rejection, etc.) for the operation order of the avatar included in the corresponding plurality of input sub-robotics videos as being in a correct state or an incorrect state, label values for adjusting the order of the operation order or the sub-robotics videos including the avatar, human, robot, etc. when the operation order of the avatar is incorrect or needs adjustment, and the identification information of the terminal 100.
[0255] Also, the server 200 receives label values for the operation order of the avatar included in the corresponding plurality of sub-robotics videos transmitted from the terminal 100 as being in a correct state or an incorrect state, label values for adjusting the order of the operation order or the sub-robotics videos including the avatar when the operation order of the avatar is incorrect or needs adjustment, and the identification information of the terminal 100.
[0256] Also, the server 200 re-aligns the order of the plurality of sub-robotics videos obtained by splitting the first robotics video based on label values such as label values for the operation order of the avatar included in the received corresponding plurality of sub-robotics videos as being in a correct state or an incorrect state, and label values for adjusting the order of the operation order or the sub-robotics videos including the avatar when the operation order of the avatar is incorrect or needs adjustment.
[0257] Also, the terminal 100 transmits to the server 200 one or more additional selection label values, one or more time-series split selection label values, one or more body part-specific selection label values, label values for aligning the order of the corresponding plurality of sub-robotics videos, and the identification information of the corresponding terminal 100 at one or more other specific time points (or other specific time intervals) related to the first robotics video.
[0258] Further, in conjunction with the server 200, before or after performing additional selection labeling on the corresponding first robotics video, the terminal 100 performs additional hierarchical labeling on the corresponding one or more first robotics videos, and may also perform additional selection labeling on the corresponding first robotics video before / after performing the additional hierarchical labeling. Here, the additional hierarchical labeling (or additional hierarchical labeling) is user input feature engineering, which attaches a label (or label value) indicating the feature of the corresponding first robotics video and divides (or classifies) the corresponding first robotics video into a plurality of sub-robotics videos according to the feature.
[0259] That is, in conjunction with the server 200, for the first robotics video displayed on the corresponding terminal 100, in relation to the corresponding specific topic, with reference to (or based on) a plurality of preset label classifications, according to the input of the user of the corresponding terminal 100 (or the user's selection / touch / control), additional labels (or additional label values) at other specific time points (or other specific intervals) in the corresponding first robotics video are set (or received / input).
[0260] At this time, when the playback bar included in the view screen within the execution result screen of the application displayed on the terminal 100 is selected, or when the playback button within the corresponding view screen is selected, the terminal 100 displays (or outputs) the first robotics video in the video display area, and displays (or outputs) a comparison target video related to the first robotics video (or the corresponding robot operation video provided by the server 200 / the comparison target video corresponding to the first robotics video) in the display area of the comparison target video. At this time, the terminal 100 may synchronize the corresponding first robotics video and the comparison target video based on the meta information corresponding to the first robotics video and the comparison target video respectively, and display the synchronized first robotics video and comparison target video in the video display area and the display area of the comparison target video respectively. Here, when either one of the first robotics video displayed in the video display area within the terminal 100 and the comparison target video displayed in the display area of the comparison target video stops due to a pause function or a stop function, the terminal 100 controls the other one to also stop due to the pause function or the stop function.
[0261] In addition, the terminal 100 sets (or receives / inputs) one or more step-by-step additional labels (or additional label values) for the movement (or behavior of the object) of the object included in the first robotics video displayed in the video display area of the terminal 100 in response to the input (or selection / touch / control of the user) of the user of the corresponding terminal 100, and also at other specific time points (or other specific intervals).
[0262] That is, the terminal 100 hierarchically inputs additional hierarchical labels (or additional hierarchical label values) for the specific operations of the object, the specific methods of the specific operations, the specific steps of the specific methods, etc. for the movement (or behavior of the object) of the object included in the first robotics video displayed in the video display area in response to the input of the user at one or more other specific time points (or other specific intervals) of the first robotics video.
[0263] In this way, for the first robotics video related to the corresponding specific topic, the terminal 100 sets (or receives / inputs) one or more additional hierarchical labels (or additional hierarchical label values) respectively at one or more other specific time points (or other specific intervals) in response to the input of the user of the terminal 100 who is an expert related to the corresponding specific topic.
[0264] Also, before / after performing the additional hierarchical labeling process, the terminal 100 performs the additional selection labeling process described above.
[0265] In this way, the terminal 100 performs functions such as an additional hierarchical labeling function and an additional selection labeling function on the first robotics video.
[0266] In the embodiments of the present invention, the additional hierarchical labeling function and the additional selection labeling function are separately described, but it is not limited thereto. The terminal 100 can include the additional hierarchical labeling function in the additional selection labeling function, and can also integrate the additional hierarchical labeling and the additional selection labeling into one additional labeling function.
[0267] Also, the terminal 100 receives a second robotics video transmitted from the server 200. Here, the second robotics video is a result generated by a classification model and a prediction model for the corresponding first robotics video in the server 200, and may be an operation-related video such as an avatar, an item, or a robot generated based on the first robotics video, or a video in which the first robotics video is updated.
[0268] In addition, the terminal 100 outputs the received second robotics video to the video display area. At this time, the terminal 100 can also divide the screen of the corresponding terminal 100 and output it simultaneously in a state where the robot operation video, the comparison target video, the first robotics video, and the second robotics video are synchronized.
[0269] In addition, the terminal 100 can be provided with a second robotics video (or an updated second robotics video) that has been newly aggregated with intelligence in relation to the specific topic (or the load data) from the server 200.
[0270] In the embodiments of the present invention, the terminal (300) is described as performing functions such as a load data collection function, a hierarchical labeling function for information / videos, a selection labeling function for information / videos, a time-series division selection labeling function for information / videos, and a selection labeling function for body parts of information / videos in the form of a dedicated application. However, the present invention is not limited to this, and in addition to the dedicated application, the load data collection function, the hierarchical labeling function for information / videos, the selection labeling function for information / videos, the time-series division selection labeling function for information / videos, the selection labeling function for body parts of information / videos, etc. can also be performed through a website provided to the server 200 or the like.
[0271] The server 200 communicates with the terminal 100 and the like.
[0272] In addition, the server 200 performs procedures such as member registration for users of the terminal 100 and the like.
[0273] In addition, the server 200 registers personal information related to users of the terminal 100 and the like. At this time, the server 200 can register (or manage) the corresponding personal information and the like in a DB server (not shown).
[0274] In addition, the server 200 performs a member management function for users of the terminal 100 and the like.
[0275] In addition, the server 200 provides the terminal 100 and the like with a dedicated application and / or a website that provides functions such as a raw data collection function, a hierarchical labeling function for information / videos, a selective labeling function for information / videos, a time-series segmentation selective labeling function for information / videos, and a selective labeling function for different body parts of information / videos.
[0276] In addition, the server 200 provides a bulletin board function for announcements, events, and the like.
[0277] In addition, the server 200, in conjunction with the terminal 100 and the payment server, performs a payment function by the execution of a subscription function on the corresponding terminal 100 for functions such as a raw data collection function, a hierarchical labeling function for information / videos, a selective labeling function for information / videos, a time-series segmentation selective labeling function for information / videos, and a selective labeling function for different body parts of information / videos provided by the corresponding server 200.
[0278] When the payment function fails, the server 200 provides the terminal 100 with payment failure information (for example, including payment date, payment amount, failure information (including insufficient balance, exceeding limit, etc.)) (or information indicating that the payment has failed).
[0279] In addition, after the payment function with the terminal 100 is successfully performed, the server 200 transmits the payment function execution results provided by the payment server to the terminal 100 respectively. Here, the payment function execution results include subscription period, payment amount, payment date, and time information, etc.
[0280] In addition, the server 200 maps (or matches / links) the payment function execution results with the corresponding terminal 100 (or account information related to the corresponding terminal 100) and manages (or saves / registers) them.
[0281] In addition, by performing the subscription function, the server 200 provides various information for performing functions such as a load data collection function provided by the corresponding server 200 via the corresponding dedicated application on the terminal 100, a hierarchical labeling function for information / videos, a selection labeling function for information / videos, a time-series division selection labeling function for information / videos, a selection labeling function for body parts of information / videos, and the like.
[0282] In addition, the server 200 may further include a bus (not shown), a communication interface (not shown), etc. to provide a communication function between components of the corresponding server 200.
[0283] The bus is implemented as various types of buses such as an address bus, a data bus, and a control bus.
[0284] The communication interface supports wired / wireless Internet communication of the server 200.
[0285] In addition, when a computer program is loaded into the memory, the server 200 includes one or more instructions that cause the processor to perform the methods / functions according to various embodiments of the present invention. That is, the processor performs the one or more instructions to perform the methods / functions according to various embodiments of the present invention.
[0286] In addition, the server 200 utilizes, as data for continuous machine learning (or deep learning), load data related to a specific topic collected in advance, meta information related to the corresponding load data, comparison target video, meta information related to the corresponding comparison target video, first video, meta information related to the corresponding first video, second video, meta information related to the corresponding second video, operation-related video of an avatar and / or an item, meta information related to the corresponding operation-related video, first robotics video, meta information related to the corresponding first robotics video, second robotics video, meta information related to the corresponding second robotics video, and so on. Here, the input data set for the machine learning is divided into a training set and a test set at a preset ratio (including, for example, 7:3, 8:2, etc.) from the load data related to the specific topic collected in advance, meta information related to the corresponding load data, comparison target video, meta information related to the corresponding comparison target video, first video, meta information related to the corresponding first video, second video, meta information related to the corresponding second video, operation-related video of an avatar and / or an item, meta information related to the corresponding operation-related video, first robotics video, meta information related to the corresponding first robotics video, second robotics video, meta information related to the corresponding second robotics video, and so on, so that training and test functions can be performed. Also, the input data set for the machine learning includes load data related to a specific topic collected subsequently, meta information related to the corresponding load data, comparison target video, meta information related to the corresponding comparison target video, first video, meta information related to the corresponding first video, second video, meta information related to the corresponding second video, operation-related video of an avatar and / or an item, meta information related to the corresponding operation-related video, first robotics video, meta information related to the corresponding first robotics video, second robotics video, meta information related to the corresponding second robotics video, and so on.In addition, the output data set for the machine learning learns, in the part to be predicted, based on information collected etc., and then classifies or predicts this, and classifies labels related to the corresponding load data, the first video, the second video, the operation-related video, the first robotics video, the second robotics video, etc., and includes the first video, the second video, the first robotics video, the second robotics video, etc. generated based on the classified information.
[0287] That is, the server 200 performs a learning function for classifying label values related to corresponding information for raw data, a first video, an avatar and / or item operation-related video, a first robotics video, etc. related to a specific topic pre-collected for the classification model through pre-set learning data. At this time, the server 200 stores the corresponding information in parallel and in a distributed manner, and purifies raw data, structured data, and semi-structured data such as raw data related to a specific topic pre-collected included in the stored information, meta information related to the corresponding raw data, a comparison target video, meta information related to the corresponding comparison target video, a first video, meta information related to the corresponding first video, a second video, meta information related to the corresponding second video, an avatar and / or item operation-related video, meta information related to the corresponding operation-related video, a first robotics video, meta information related to the corresponding first robotics video, a second robotics video, meta information related to the corresponding second robotics video, etc., and performs pre-processing including classification as meta data, performs analysis including data mining on the pre-processed data, and can construct big data by proceeding with learning, training, and testing based on at least one type of machine learning. At this time, at least one type of machine learning may consist of any one or at least one combination of supervised learning, semi-supervised learning, unsupervised learning, reinforcement learning, and deep reinforcement learning.
[0288] In addition, the server 200 performs a learning function for generating new videos (for example, including the first video, the second video, etc.) related to the corresponding information with respect to the classification values, raw data, meta information related to the corresponding raw data, comparison target videos, meta information related to the corresponding comparison target videos, the first video, meta information related to the corresponding first video, the second video, meta information related to the corresponding second video, operation-related videos of the avatar and / or item, meta information related to the corresponding operation-related videos, the first robotics video, meta information related to the corresponding first robotics video, the second robotics video, meta information related to the corresponding second robotics video, etc. classified by the classification model in relation to a specific topic collected in advance for the prediction model through the preset learning data. At this time, the server 200 stores the corresponding information in parallel and in a distributed manner, purifies the classification values, raw data, meta information related to the corresponding raw data, comparison target videos, meta information related to the corresponding comparison target videos, the first video, meta information related to the corresponding first video, the second video, meta information related to the corresponding second video, operation-related videos of the avatar and / or item, meta information related to the corresponding operation-related videos, the first robotics video, meta information related to the corresponding first robotics video, the second robotics video, meta information related to the corresponding second robotics video, etc. included in the stored information into unstructured data, structured data, and anti-structured data, performs preprocessing including classification as meta data, performs analysis including data mining on the preprocessed data, and can construct big data by proceeding with learning, training, and testing based on at least one type of machine learning. At this time, at least one type of machine learning may consist of any one or at least one combination of supervised learning, semi-supervised learning, unsupervised learning, reinforcement learning, and deep reinforcement learning.
[0289] In this way, the server 200 performs a learning function for the classification model, the prediction model, etc. in the form of neural networks through the learning data, etc.
[0290] In addition, the server 200 uses a generative neural network algorithm, a tracking neural network, etc. Here, the tracking neural network may be a neural network algorithm that can measure and structure data by taking the counterpart values of the xyz coordinates of the video information of an object as a 4-dimensional vector product while being a model into which sequential input is input.
[0291] In one embodiment of the present invention, GNN (Graph Neural Network), GAN (Generative Adversarial Network), etc. are used as the generative neural network algorithm and the tracking neural network. There may be a combination of GAN and GNN as an artificial intelligence algorithm, there may be a single application of GNN excluding GAN, and there may be a single application of GAN excluding GNN. When using GAN alone, when obtaining the predicted values of attributes and target attributes without using the "GNN regression model type 1" and "GNN regression model type 2", deep learning and related rules are used. GAN enhances the expression of still images and videos, the naturalness and delicacy of image quality, etc. Infer using related rules of operation patterns to predict the next operation.
[0292] The first basic video information which is raw data becomes the first attribute 1224 and the first target attribute 1225 clustered by the first hierarchical labeling 1210.
[0293] Referring to FIGS. 3 to 5, when the user performs the first hierarchical labeling 1210 on a plurality of basic videos (or raw data / basic video information) in the annotation step, the basic video information is hierarchically clustered. This is used as the first hierarchical cluster, and the basic video indicates the video in which the basic video information is output on the view screen of the terminal 100.
[0294] In addition, the server 200 receives one or more load data transmitted from the terminal 100, meta information related to the corresponding load data, comparison target videos, meta information related to the corresponding comparison target videos, identification information of the terminal 100, and the like.
[0295] At this time, when the comparison target video related to the load data is not transmitted from the terminal 100, the server 200, based on one or more load data related to the received specific topic, meta information related to the corresponding load data, etc., confirms (or searches) the comparison target video related to the load data among the plurality of comparison target videos managed by the corresponding server 200, and provides the confirmed corresponding comparison target video, meta information related to the corresponding comparison target video, etc. to the terminal 100.
[0296] In addition, the server 200 performs selection labeling on the received one or more load data. Here, the selection labeling (or selection labeling) refers to a labeling method of setting (or attaching) a label (or label value) for the presence or absence of an error (or abnormality) at a specific time point (or specific section) of the load data. At this time, for the time point (or section) of the load data for which a label (or label value) has not been set by the selection labeling, a preset default label value (for example, an approval label) may be set.
[0297] That is, the server 200, in conjunction with the terminal 100, sets (or receives / inputs) the label (or label value) at a specific time point (or specific section) of the corresponding load data in response to the input of the user of the corresponding terminal 100 (or the user's selection / touch / control) for the load data displayed on the corresponding terminal 100.
[0298] In addition, the server 200 receives one or more selection label values at one or more specific time points (or specific sections) related to the load data transmitted from the terminal 100, meta information of the corresponding load data, identification information of the corresponding terminal 100, and the like.
[0299] In an embodiment of the present invention, in the terminal 100, mainly described is setting (or receiving / inputting) one or more selection label values at one or more specific time points (or specific intervals) among the corresponding load data in response to a user's input. However, the present invention is not limited thereto. The server 200 performs a video analysis function on the corresponding load data and a comparison target video related to the corresponding load data, and based on the result of performing the video analysis function, automatically sets one or more selection label values for the corresponding load data at one or more specific time points (or specific intervals).
[0300] Also, in the server 200, when one or more selection label values are set for the corresponding load data at one or more specific time points (or specific intervals), the server 200 provides information regarding the one or more selection label values at one or more specific time points (or specific intervals) related to the set corresponding load data to the terminal 100. In the corresponding terminal 100, information regarding the one or more selection label values at one or more specific time points (or specific intervals) related to the corresponding load data set by the server 200 is displayed, and according to the input of the user of the corresponding terminal 100, it may be configured to determine the presence or absence of final approval for the one or more selection label values at the corresponding one or more specific time points (or specific intervals).
[0301] At this time, before or after performing selection labeling on the corresponding one or more load data, the server 200 performs hierarchical labeling on the corresponding one or more load data in conjunction with the terminal 100, and it is also possible to perform selection labeling on the corresponding one or more load data before / after performing hierarchical labeling. Here, the hierarchical labeling (or hierarchical leveling) refers to a labeling method that, as user input feature engineering, attaches a label (or label value) indicating a feature for the corresponding load data and divides (or classifies) the corresponding load data into a plurality of sub-load data according to the feature.
[0302] That is, the server 200, in conjunction with the terminal 100, refers to (or is based on) a plurality of preset label classifications related to the corresponding specific topic for the load data displayed on the corresponding terminal 100, and according to the input of the user of the corresponding terminal 100 (or the user's selection / touch / control), sets (or receives / inputs) the label (or label value) at another specific time point (or another specific section) of the corresponding load data.
[0303] In addition, the server 200 divides the load data into a plurality of sub-load data.
[0304] In the embodiments of the present invention, mainly described is that in the terminal 100, one or more hierarchical label values are set (or received / input) at one or more other specific time points (or other specific sections) of the corresponding load data according to the input of the user. However, it is not limited thereto. The server 200 performs a video analysis function on the corresponding load data and the comparison target video related to the corresponding load data, and based on the execution result of the video analysis function, may automatically set one or more hierarchical label values for the corresponding load data at one or more other specific time points (or other specific sections), respectively.
[0305] In addition, when the server 200 sets one or more hierarchical label values for the corresponding load data at one or more other specific time points (or other specific sections), the server 200 provides the terminal 100 with information on one or more hierarchical label values at one or more other specific time points (or other specific sections) related to the set corresponding load data. In the corresponding terminal 100, information on one or more hierarchical label values at one or more other specific time points (or other specific sections) related to the corresponding load data set by the server 200 is displayed, and according to the input of the user of the corresponding terminal 100, it may be configured to determine the presence or absence of final approval for one or more hierarchical label values at one or more other specific time points (or other specific sections).
[0306] In addition, the server 200 calls a library related to input feature engineering and converts the basic video information (or raw data) into an input feature vector. Hierarchical labeling by the user divides the basic video information into data unit 3 or data unit 4, and trains a predictive model with supervision so that the attribute values of data unit 3 or data unit 4 become composite input features. The composite input feature indicates that the basic video information combined with, for example, point cloud, RGB, JPG, video information, voxels (or 3D images), vector format, etc. has been converted into an input feature.
[0307] In addition, the hierarchical labeling by the user may be omitted in the process where the server 200 calls a library related to input feature engineering to convert the basic video information into an input feature vector.
[0308] The first hierarchical cluster 1201 may be generated by the server 200 itself. When the user does not perform some or all of the first hierarchical labeling, second hierarchical labeling, etc. by the user, it can be said that the hierarchical clustering labeling is performed by the server 200 for the artificial intelligence to obtain the input features by itself.
[0309] In one embodiment of the present invention, the step of receiving hierarchical labeling information for receiving hierarchical clustering labeling information is omitted. Input feature engineering by the user, such as the first hierarchical labeling, second hierarchical labeling, third hierarchical labeling, etc., is omitted, and the server 200 obtains the input features by itself. Input feature engineering by the user, such as repeated hierarchical labeling like the first hierarchical labeling, second hierarchical labeling, third hierarchical labeling, etc., is omitted, and the server 200 obtains the input features by itself.
[0310] In FIG. 12, the first hierarchical cluster 1201 indicates that the first basic video information is clustered by the first hierarchical labeling 1210. The first hierarchical cluster 1201 becomes the hierarchical cluster 700 based on the data unit 3 in FIG. 7 or the hierarchical cluster 900 based on the data unit 4 in FIG. 9.
[0311] In one embodiment of the present invention, the hierarchical cluster includes a method in which the server 200 requests input features by itself.
[0312] The second hierarchical labeling in FIG. 12 advances the hierarchical clustering labeling for the first video information output on the view screen of the terminal 100, so that the user refers to the label classifications in [Table 1] to [Table 11] above to input the hierarchical clustering label value for the first video information.
[0313] In one embodiment of the present invention, the user (or the terminal 100 / server 200) does not perform labeling for each specific step and / or each detailed operation step. The video may be divided into the data unit 3, the data unit 4, or the data unit 5 by the server 200.
[0314] Further, the server 200 performs machine learning on the artificial intelligence platform based on information about the selected labeled load data, etc., and generates (or confirms) a classification value for the corresponding load data based on the result of the machine learning. Here, the classification value for the corresponding load data (or the classification value of the corresponding load data / the classification value of the selected labeled load data / the classification value of the hierarchically labeled load data) may be a value obtained by classifying selection labeling values, hierarchical labeling values, etc. by the same item.
[0315] That is, the server 200 performs machine learning (or artificial intelligence / deep learning) using information about the selected labeled load data as input values of a preset classification model, and generates (or confirms) a classification value for the corresponding load data based on the result of the machine learning (or the result of artificial intelligence / the result of deep learning).
[0316] In various embodiments, in the labeling step, the labeling for classifying the operations of an avatar, a human, a robot, etc. as approved (ACCEPT) or rejected (REJECT) is carried out in the form of supervised learning, which corresponds to a classification model. The approval (ACCEPT) and rejection (REJECT) binary classification can be used as a commonly used binary classification model, and when expressing the success or failure of a surgery and an operation on a 5-step scale, it can be implemented as a multiple classification model from which the probability value of each class is derived.
[0317] In various embodiments, the video information can be labeled by a dichotomy of selecting approval (ACCEPT) or rejection (REJECT) on the user interface of the execution result screen (or view screen) of the application of the terminal 100, but it can also be labeled in 3 steps by classifying it into approval (ACCEPT), normal (NORMAL), and rejection (REJECT). It is also possible to further divide and label the steps of correct and incorrect operations into 5-step or 6-step labels according to the degree. When the label subdivision is large, such as 5 steps or 6 steps, the good scores are scored from 5 points to 1 point. If the good score is above a certain value (4 points or more), it is regarded as approved (ACCEPT), and if the incorrect score is below a certain value (2 points or less), it is regarded as rejected (REJECT) for classification. 3 points are classified as normal (NORMAL).
[0318] Further, the server 200 performs machine learning (or artificial intelligence / deep learning) using, as input values, the classification value for the generated corresponding load data (or the classification value of the corresponding load data), information regarding the selected and labeled load data, the corresponding load data, meta information related to the corresponding load data, the comparison target video, meta information related to the corresponding comparison target video, etc., and generates a first video corresponding to the corresponding load data based on the result of the machine learning (or the result of artificial intelligence / the result of deep learning). At this time, the first video may be an operation-related video such as an avatar, an item, or a robot generated based on the load data, a video in which the load data is updated (for example, a video in which the actions / behavior / conduct of a human / humans included in the load data are updated), etc.
[0319] That is, the server 200 performs machine learning (or artificial intelligence / deep learning) using, as input values of a preset prediction model, the classification value for the generated corresponding load data (or the classification value of the corresponding load data), information regarding the selected and labeled load data, the corresponding load data, meta information related to the corresponding load data, the comparison target video, meta information related to the corresponding comparison target video, etc., and generates a first video related to the corresponding load data based on the result of the machine learning (or the result of artificial intelligence / the result of deep learning).
[0320] Further, the server 200 transmits (or provides) the generated first video to the terminal 100.
[0321] Referring to FIG. 13, the structure of the GNN is as follows.
[0322] Objects in videos and still images are represented by nodes (x1~x4, z1~z4). Each object is related to each other, and there is a time-series movement pattern in which the relationship has a mutual influence. The input layer 1301 and the output layer 1303 have multiple layers overlapping, and there is a hidden layer (or hidden layer) 1302 between the input layer 1301 and the output layer 1303. When an input is entered, the next output is predicted.
[0323] Existing GANs use a 3D voxel method. When voxelizing space in 3D, X * , Y * , Z * , the information in the 4 dimensions reaches several hundred megabytes, so there are problems that require a very large amount of hardware, GPUs, and memory resources, and a very long training time is required. Due to such problems, recently, the point cloud method is mainly used. The point cloud method can measure the physical space using Lidar, etc., and it is possible to physically measure and structure the corresponding values of xyz coordinates, which has an effective aspect compared to the 3D voxel method. However, since the points are not aligned and not shaped, there is a disadvantage that only a very small part of the characteristics of the object is expressed in the artificial intelligence, and it is necessary to express the aligned information for relative values and features.
[0324] The GAN according to an embodiment of the present invention further expresses characteristics of joints (for example, can only bend inside), angles, distances, landmark points, etc. in the connection between points in order to express information related to the characteristics of videos, poses, and movements within the features.
[0325] In one embodiment of the present invention, the point cloud may be embodied in other data structures within a range that does not deviate from the essential features.
[0326] That is, the GAN according to the embodiment of the present invention can represent the characteristics of joints and structures in a vector product and has a replaced data structure from the point cloud form to the GNN form.
[0327] Also, in the utilization of GAN for 3D space, 3D motion, body shape, and movement, the information of the space is generated by an object in which a plurality of points that are not points are combined, and this is processed by structuring the data in the form of a GNN.
[0328] When processing in the form of a GNN, additional meta information is used as another feature of the input value. When the form of the feature of the meta information is different, it cannot be simply changed, so it is used by separating and merging the layers.
[0329] The meta information merged and used as other input values includes user information and item information.
[0330] The meta information is used as complementary information for the supervised label and as conditional information in the training of the unsupervised GAN.
[0331] The corresponding meta information enables the GAN to remember the degree of similarity during various visual trainings. When the subsequent specific attribute information changes, it is utilized as input value information that is useful when the visual information is variably and artificially intervened accordingly. In one embodiment of the present invention, when increasing the numerical value of muscle mass or decreasing the age, the shape of the generated virtual avatar can change according to the corresponding value of the meta information.
[0332] The GNN can represent an artificial neural network structure implemented in a manner that derives similarities and feature points between modeling data using the modeling data modeled based on the mapped data between specific parameters. Here, other things can be used in addition to the algorithms covered, and it is not limited to the mentioned algorithms.
[0333] In one embodiment of the present invention, the user information includes the form and hue of the face and body, age, gender, hair, race, fat level, muscle mass level, and other various category information, numeric information, and other user attribute information. The item information includes the brand, producer ID, advertiser ID, NFT ID, product group ID, and other item attribute information. When utilized in a digital device, it is information such as the name of each part, blood type, age, gender, disease type, and progression status.
[0334] Referring to FIG. 14, the server 200 modifies 1401 the condition meta information of the Conditional GAN and modifies 1401 the body shape characteristic information such as the avatar becoming slimmer or being a muscle man. Referring to FIG. 14, the meta information can be modified 1401 in various games.
[0335] In one embodiment of the present invention, the server 200 generates (or manages) an avatar that adjusts a dancing performance, virtual surgery, virtual soccer game, virtual fighter, etc.
[0336] In one embodiment of the present invention, the digital caddy may be an external object that can be alternated with prostheses, implants, etc. during dental surgery. It can be alternated and simulated before surgery. In shaping, it is a simulation after shaping. In general surgery, it can be utilized as a physical coupling simulation application with 3D three-dimensional size and structure. Thus, the characteristics of the pre-learned object (for example, including the opening and closing of various medical equipment, the hands and feet of the doctor cannot leave the body and can bend inward, medical equipment and instruments can leave the digital caddy and approach, etc.) are utilized as training features.
[0337] Thus, the characteristics of the pre-learned object (for example, including that the dental blade (bar) of the dental handpiece can rotate, the tissue is opened by the surgical blade, the tooth can be removed from the gum, the human organ can be replaced, etc.) are utilized as the training features.
[0338] Thus, the characteristics of the pre-learned object (for example, including that the tire of the car can rotate, the front door of the house can be opened, the hands and feet cannot leave the body and can bend inward, the hat can leave the head and can be put on, etc.) are utilized as the training features.
[0339] Thus, the server 200 modifies 1401 the conditional meta information of the conditional GAN and modifies the mutation of the digital caddy, which is a kind of avatar, and various disease information according to the case.
[0340] As various embodiments, the terminal 100 may be a VR simulator in various forms. For a VR simulator provided with visual rendering by a GAN and / or GNN prediction model, haptic rendering is simultaneously provided. A visual set device and various forms of haptic devices are connected to the VR simulator. The types according to the form of the VR simulator are as follows. That is, the VR simulator includes a tooth extraction VR simulator, a surgical VR simulator, a VEHICLE VR simulator, a VR treadmill, and the like. The form of the VR simulator is not limited to this.
[0341] In one embodiment of the present invention, the equipment of the tooth extraction VR simulator requires an HMD, a haptic device, a foot pedal system used in a dental chair (including, for example, Arduino, Raspberry Pi, etc.). Use 3D printing to create a digital cad in virtual reality and create an artificial cad with HD haptics. Use VR and 3D simulators to perform virtual dental treatment and surgery.
[0342] In one embodiment of the present invention, the surgical VR simulator is as follows. A 3D model of the patient's pathological valve is created, and the 3D patient coordinate system is aligned with the coordinate system of the patient placed on the operating table based on the position and state of the pathological valve and video information, so as to predict the position of the invisible pathological valve and perform the surgery.
[0343] In one embodiment of the present invention, for various forms of VR simulators (examples of VEHICLE: submarine, tank, drone, fighter, etc.), data can be obtained on the method of driving an avatar, a human, a robot, etc. using the control device of the VEHICLE-type VR simulator to generate an avatar.
[0344] The VR vehicle simulator simulates using the arms, legs, and other body parts of the operator's own avatar. The coordinate system is synchronized according to rules from start to end in the metabus world. In order to implement a high level of visual rendering, it must be equipped with a lidar, infrared tracking, and motion tracking, and also have a motion data alignment algorithm for the human body, and an alignment algorithm for the position of the simulator in the metabus world is also necessary.
[0345] For example, the visual data (basic video information) obtained by virtual airplane operation becomes a dataset of the initial model of the guidance and / or inference algorithm in FIG. 12.
[0346] The guidance and / or inference algorithm 1200 in FIG. 12 is the sum of the partial guidance and / or inference algorithms (the first and second guidance and / or inference algorithms) in FIG. 15.
[0347] The visual data (basic video information) obtained by virtual airplane operation becomes the basic data that enables artificial intelligence to operate the flight simulator. In virtual flight operation of artificial intelligence, when many errors and mistakes occur, the user (airplane pilot) proceeds with selection labeling through the user interface of the execution result screen (or view screen) of the app on the terminal 100.
[0348] In one embodiment of the present invention, an avatar control system (using a HEAD MOUNTED DISPLAY) using a VR treadmill requires the following technologies.
[0349] An integrated algorithm for the movement, actions, infinite walking, rotation, etc. of the user and the avatar, a posture control system, a motion and movement control system using the VIVE Tracker, an integrated algorithm for infinite walking and human motion data using a rider and infrared tracking (utilizing the pressure value of the shoes and the infrared sensor value), a VR treadmill body designed to enable almost all movements of the human body, a reaction technology based on the coordinate reference and environmental variation of the metaverse world, a motion data synchronization and dedicated server, a synchronization system that enables the user's network play, etc. are required.
[0350] Among the total of K2 to K6 clusters in FIGS. 7 to 11, a regression model using a GNN for the coordinate values of the static image of the attribute and various visual data belonging to a specific cluster is defined as the "Type 1 GNN Regression Model", and a regression model using a GNN for the coordinate values of the video and various visual data for the target attribute is defined as the "Type 2 GNN Regression Model".
[0351] Referring to FIG. 13, a model for predicting the relative video information and state values of a specific viewpoint with respect to the action behavior of the avatar is structured in the form of a GNN, and the model that predicts each of these numerical values is defined as the "Type 1 and Type 2 GNN Regression Models".
[0352] When using GAN alone, the First Related Rule Type 1 1214 and the First Related Rule Type 2 1215 predict the Second Attribute 1226 and the Second Target Attribute 1227. The Related Rule Type 1 and Type 2 are models that infer static images and videos respectively using related rules and deep learning (a model that takes an input in the Sequential form excluding the GNN regression model) without using a GNN, and are models in the same form as the Type 1 and Type 2 GNN Regression Models in FIG. 13 except that they do not use the GNN form of structuring.
[0353] In one embodiment of the present invention, the deep learning used in the tracking neural network (a model in which the input in the Sequential form except for the GNN regression model is input, and the x, y, and z coordinates of the object are tracked) is a deep neural network.
[0354] "Type 1 and Type 2 GNN regression models" or "Type 1 and Type 2 related rules" are models that use the sliding window technique and receive input in the Sequential form. "Type 1 and Type 2 GNN regression models" or "Type 1 and Type 2 related rules" are the "GAN and / or GNN prediction model 1605" in FIG. 16.
[0355] Referring to FIG. 16, in the user interface 1603 of the terminal 100 connected to the visual set device 1602, the user performs selection labeling 1604, and the labeled visual data is used as the GAN and / or GNN prediction model 1605. The GAN and / or GNN prediction model 1605 transmits the visual data to the simulation engine 1606 in order to generate or output the actions of the avatar. In FIG. 16, the visual data is transmitted in the order of the simulation engine 1606, the graphics engine 1607, the display device 1608, and the control algorithm 1609, and is output via the user interface 1603. The execution result screen (or view screen) of the application of the terminal 100 is the one in which the user interface 1603 is embodied on the screen from the terminal 100.
[0356] The GAN and / or GNN prediction model 1605 in FIG. 16 includes an interface API process.
[0357] In various embodiments, examples of interface APIs are as follows. Data received by an IoT Edge device (including, for example, Arduino, Raspberry Pi, etc.) may be the input data itself or the output of the result of artificial intelligence inference driven on the Edge. An artificial intelligence model created in Python, etc., can be converted for use on an IoT Edge device by an open source library such as ONNX. As a result, the output result data and input data inferred at the primary level on the Edge are re-inferred by a more complex collective intelligence model through a Server API call.
[0358] The digital unit means a video unit divided by the interaction between artificial intelligence and the user (including, for example, time-series segmentation selection labeling, selection labeling by body part, etc.).
[0359] Classify the first attribute 1224 selected and labeled 1604 for each of K2 or K4 clusters from the hierarchical clusters of FIG. 7 or FIG. 9, and induce and / or infer (ai inference) the first GNN regression model type 1204 or the first association rule type 1214.
[0360] With the first attribute 1224 and the first target attribute 1225 clustered for each of K2 or K4 clusters from the hierarchical clusters of FIG. 7 or FIG. 9, induce and / or infer (ai inference) the first GNN regression model type 2 1205 or the first association rule type 2 1215.
[0361] When a time series sequence of still image information (data units 1 and 2 or the first attribute 1224) is input into the first type 1204 of the GNN regression model or the first type 1214 of the associated rule, the first type 1204 of the GNN regression model or the first type 1214 of the associated rule returns a time series sequence of the second attribute 1226 to the execution result screen (or view screen) of the application on the terminal 100. The second attribute 1226 is the predicted values 1206 and 1216 of the first type 1204 of the GNN regression model or the first type 1214 of the associated rule, and is a feature vector representation for the still image information at the k-th, L-th, or f-th step of the motion video.
[0362] When a time series sequence of the "second attribute 1226" is input into the second type 1205 of the GNN regression model or the second type 1215 of the associated rule, the second type 1205 of the GNN regression model or the second type 1215 of the associated rule generates and outputs a "second target attribute 1527", which is the predicted values 1207 and 1217 of the second type 1205 of the GNN regression model or the second type 1215 of the associated rule, to the execution result screen (or view screen) of the application on the terminal 100. The second target attribute 1227 is a feature vector representation for the video information at the k-th, L-th, or f-th step of the motion video.
[0363] Referring to FIGS. 12 and 15, the first induction and / or inference (ai inference) algorithm 1502 is as follows. The data of the first hierarchical cluster 1201 is first labeled 1202 and the first classification model 1203 is induced and / or inferred (ai inference). The classified first attribute 1224 and the first target attribute 1225 are used for the induction and / or inference of the first GAN and / or GNN prediction model 1508.
[0364] In addition, the server 200 performs additional selection labeling on the first video. Here, the additional selection labeling (or additional selection relabeling) refers to a labeling method of setting (or attaching) a label (or label value) for the presence or absence of an error (or abnormality) at another specific time point (or another specific section) of the first video. At this time, for the time point (or section) of the first video for which no label (or label value) is set by the additional selection labeling, a preset default label value (for example, an approval label) may be set.
[0365] That is, the server 200, in conjunction with the terminal 100, sets (or receives / inputs) a label (or label value) at another specific time point (or another specific section) of the corresponding first video in response to the input of the user of the corresponding terminal 100 (or the user's selection / touch / control) for the first video displayed on the corresponding terminal 100.
[0366] In addition, the server (00) receives one or more additional selection label values, one or more time-series division selection label values, one or more selection label values by body part, label values for sorting the order of a plurality of sub-videos, identification information of the corresponding terminal 100, etc. at one or more other specific time points (or another specific section) related to the first video transmitted from the terminal 100.
[0367] In the embodiments of the present invention, mainly described is that in the terminal 100, one or more additional selection label values are set (or received / input) at one or more other specific time points (or another specific section) of the corresponding first video in response to the input of the user. However, the present invention is not limited thereto. The server 200 performs a video analysis function on the corresponding first video and a comparison target video related to the corresponding first video, and based on the result of performing the video analysis function, automatically sets one or more additional selection label values at one or more other specific time points (or another specific section) for the corresponding first video.
[0368] Also, when one or more additional selection label values are set for the corresponding first video at one or more other specific time points (or other specific intervals) in the server 200, the server 200 provides information regarding the one or more additional selection label values at one or more other specific time points (or other specific intervals) related to the set corresponding first video to the terminal 100. In the corresponding terminal 100, information regarding the one or more additional selection label values at one or more other specific time points (or other specific intervals) related to the corresponding first video set by the server 200 is displayed, and the presence or absence of final approval is determined for the one or more additional selection label values at one or more other specific time points (or other specific intervals) according to the input of the user of the corresponding terminal 100.
[0369] At this time, before or after performing additional selection labeling for the corresponding first video, the server 200 performs additional hierarchical labeling for the corresponding one or more first videos in conjunction with the terminal 100. Before / after performing the additional hierarchical labeling, it is also possible to perform additional selection labeling for the corresponding first video. Here, the additional hierarchical labeling (or additional hierarchical leveling) is user input feature engineering, which indicates a labeling method of attaching a label (or label value) indicating the feature of the corresponding first video and dividing (or classifying) the corresponding first video into a plurality of sub-videos according to the feature.
[0370] That is, the server 200, in conjunction with the terminal 100, refers to (or based on) a plurality of preset label classifications related to the corresponding specific topic for the first video displayed on the corresponding terminal 100, and sets (or receives / inputs) additional labels (or additional label values) at one or more other specific time points (or other specific intervals) in the corresponding first video according to the input (or user selection / touch / control) of the user of the corresponding terminal 100.
[0371] Also, the server 200 divides the first video into a plurality of sub-videos.
[0372] In an embodiment of the present invention, in the terminal 100, mainly described is setting (or receiving / inputting) one or more additional hierarchical labels (or additional hierarchical label values) at one or more other specific time points (or other specific intervals) among the corresponding first videos according to a user's input. However, the present invention is not limited thereto. The server 200 performs a video analysis function on the corresponding first video and a comparison target video related to the corresponding first video, and based on the result of performing the video analysis function, automatically sets one or more additional hierarchical label values for the corresponding first video at one or more other specific time points (or other specific intervals).
[0373] Also, in the server 200, when setting one or more additional hierarchical label values for the corresponding first video at one or more other specific time points (or other specific intervals), the server 200 provides information on the one or more additional hierarchical label values at the one or more other specific time points (or other specific intervals) related to the set corresponding first video to the terminal 100. In the corresponding terminal 100, information on the one or more additional hierarchical label values at the one or more other specific time points (or other specific intervals) related to the corresponding first video set by the server 200 is displayed, and according to the input of the user of the corresponding terminal 100, it may be configured to determine whether to finally approve the one or more additional hierarchical label values at the one or more other specific time points (or other specific intervals).
[0374] When performing a second - level labeling 1220 on the first video information 1503, which is the predicted value of the first GNN and / or GAN prediction model 1508 in FIG. 15, a second - level cluster 1507 is obtained. At the same time, a first - level labeling is performed on the second basic video information 1506, and the generated cluster is included in the second - level cluster 1507.
[0375] In one embodiment of the present invention, the information reception step of the second hierarchical labeling 1220 for receiving hierarchical clustering labeling information is omitted. Input feature engineering by the user is omitted and can be generated by the server 200 itself.
[0376] In an embodiment of the present invention, the second hierarchical labeling 1220 may be implemented included in the second selection labeling 1208.
[0377] Also, the server 200 performs other machine learning on the artificial intelligence platform based on information about the first video with the additional selection labeling, etc., and generates (or confirms) a classification value for the corresponding first video based on the result of the other machine learning. Here, the classification value for the corresponding first video (or the classification value of the corresponding first video) may be a value obtained by classifying additional selection labeling values, additional hierarchical labeling values, etc. by the same item.
[0378] That is, the server 200 performs other machine learning (or other artificial intelligence / other deep learning) using information about the first video with the additional selection labeling, etc. as input values of the preset classification model, and generates (or confirms) a classification value for the corresponding first video based on the result of the other machine learning (or the result of the other artificial intelligence / the result of the other deep learning).
[0379] The server 200 includes an induction and / or inference step of a second classification model 1209 for inducing and / or inferring a classification model for the basic videos (or raw data) input by K7 (several times) * K8 (several people) users.
[0380] The value predicted by the first GAN and / or GNN prediction model 1508 is the first video information 1503. The still image information of the first video information 1503 is the second attribute 1226, and the video information is the second target attribute 1227.
[0381] The second classification model 1209 classifies the "second attribute 1226 and the second target attribute 1227" that belong to a specific cluster, which is one of the second hierarchical clusters 1507. When the user inputs a hierarchical clustering label value for the second attribute 1226 and the second target attribute 1227 that belong to the specific cluster and performs "selection labeling 1604", a classification model for the labeled data is induced and / or inferred. From the classification model, the first attribute 1224 and the first target attribute 1225 of the second basic video information 1505 are learned as one model.
[0382] In an embodiment of the present invention, when the step of receiving the second hierarchical labeling information of the first video information based on the first basic video information by the user who receives the hierarchical clustering labeling information and the step of receiving the first hierarchical labeling information based on the second basic video information are omitted, the hierarchical cluster is generated by the server 200 itself.
[0383] Further, the server 200 uses, as input values, the classification value for the generated corresponding first video (or the classification value of the corresponding first video), information regarding the additionally selected and labeled first video, the corresponding first video, meta information related to the corresponding first video, the comparison target video, meta information related to the corresponding comparison target video, etc., and performs other machine learning (or other artificial intelligence / other deep learning). Based on the result of the other machine learning (or the result of the other artificial intelligence / the result of the other deep learning), a second video corresponding to the corresponding first video is generated. At this time, the second video may be an avatar, an item, an operation-related video such as a robot, etc. generated based on the first video, or a video in which the first video is updated.
[0384] That is, the server 200 uses, as input values for the preset prediction model, the classification value for the generated corresponding first video (or the classification value of the corresponding first video), information regarding the additionally selected and labeled first video, the corresponding first video, meta information related to the corresponding first video, the comparison target video, meta information related to the corresponding comparison target video, etc., to perform other machine learning (or other artificial intelligence / other deep learning), and based on the result of the other machine learning (or the result of the other artificial intelligence / the result of the other deep learning), generates a second video related to the corresponding first video.
[0385] Further, the server 200 transmits (or provides) the generated second video to the terminal 100.
[0386] "First basic video information, second basic video information, …" is input into the induction and / or inference (AI inference) algorithm 1200 in FIG. 12, and is actual visual data continuously collected via the visual set device 1602.
[0387] The second induction and / or inference algorithm 1504 is as follows. The predicted value of the first GAN and / or GNN prediction model 1508 becomes the second hierarchical cluster 1507 by the second hierarchical labeling 1220. The data of the second hierarchical cluster 1507 is secondarily selected and labeled 1208, and the second classification model 1209 is induced and / or inferred. The classified second attribute 1226 and the second target attribute are used for the induction and / or inference of the second GAN and / or GNN prediction model 1509. Hereinafter, the induction and / or inference of the algorithm is repeated.
[0388] The first GAN and / or GNN prediction model 1508 is the first GNN regression model type 1 1204 and the first GNN regression model type 2 1205, or the first association rule type 1 1214 and the first association rule type 2 1215.
[0389] The second GAN and / or GNN prediction model 1509 is a model induced and / or inferred by the second hierarchical cluster 1507 of the first video information 1503 based on the first basic video information 1501 being secondarily labeled 1208 and classified into the second classification model 1209, with the second attribute 1226 and the second target attribute 1227. The second hierarchical cluster 1507 is a cluster of the second attribute 1226 and the second target attribute 1227. Also, the first attribute 1224 and the first target attribute 1225 based on the second basic video information 1505 are also used for the induction and / or inference of the second GAN and / or GNN prediction model 1509.
[0390] Referring to FIG. 15, the first video information and the second basic video information are learned as a single model from the second induction and / or inference algorithm 1504, and the second video information 1505 is generated for each cluster. The second attribute 1226 and the second target attribute 1227 of the first video information 1503 based on the first basic video information 1501, and the first attribute 1224 and the first target attribute 1225 based on the second basic video information 1505 are labeled and learned as a single model.
[0391] In one embodiment of the present invention, the model for predicting video information that is the result of good operation is the first type 2 GNN regression model 1205 or the first type 2 association rule 1215.
[0392] In various embodiments, the "first type 2 GNN regression model 1205" uses association rules. The video information (target attribute) including the object pattern and physical attribute value is predicted by association rules from the digital unit still image information (attribute) belonging to a specific cluster in FIGS. 8, 10, and 11.
[0393] In one embodiment of the present invention, the type 2 GNN regression model is divided into a method using reverse association rules, a method using forward association rules, and a method using both forward and reverse association rules.
[0394] The second basic video information 1505 and the first video information 1503 are learned as the same model (single model). The first video information 1503 is the labeling data of the first basic video information 1501.
[0395] In FIG. 15, from the model perspective, the second basic video information 1505 and the first video information 1503, which is the predicted value of the first induction and / or inference algorithm, may have different advantages and disadvantages in terms of accuracy and refinement, but are used as the training data for the second induction and / or inference algorithm 1504.
[0396] The first video information 1503 that has been initially labeled is secondarily labeled by the labeling, and this process is continuously repeated. This process becomes such that the previously labeled data (the first video information, 1503) and other new data (the second basic video information, 1505) are repeatedly performed together. The data and / or similar label values that have been learned once also continue to appear in each repeated learning (epoch), and a process that goes through several experiments is required. Each epoch is divided into learning operation units (mini batch size) by the total number of cumulative unit labels (batch size) to perform various experiments. In the corresponding process, the label values of collective wisdom are selected and averaged and reflected in the model.
[0397] The "first induction and / or inference algorithm 1502, second induction and / or inference algorithm 1504,..." in FIG. 15 is the induction and / or inference algorithm 1200 in FIG. 12, which represents the overall algorithm as the sum of partial algorithms.
[0398] In various embodiments, it is possible to easily fabricate and initialize a digital cadaver through 3D printing simulation in a virtual space, and its use proceeds with a virtual surgery audition game to reduce the constraints of the virtual space. The surgical patterns collected through actual medical institutions, virtual surgery auditions, etc. are clustered and patterned to create an initial model of artificial intelligence. For the degree of refinement and success of the surgery, the patterns of verified specialists are separately extracted and used for supervised learning. Each surgical medical artificial intelligence initially modeled in the above manner performs virtual surgery (including, for example, procedures and treatments) on digital and artificial cadavers using a VR simulator. Then, it is gamified in a way that compensates the doctor for labeling the virtual surgery information performed by the medical artificial intelligence. The medical artificial intelligence surgery labeling is refined by a method in which the doctor directly performs surgery in the virtual space or reviews and corrects the surgery performed by the learned artificial intelligence. The reinforcement of such labeling behavior is gamified through compensation.
[0399] In one embodiment of the present invention, the sliding window is as follows. The unit of the size of each window (windows size) is classified for video information. Assuming that a 50-second video is divided into 5 parts of 10 seconds each in the above manner, if the inputs come in the order of A, B, Z, A, B, the next being Z is predicted according to the association rule.
[0400] In one embodiment of the present invention, well-known deep learning algorithms such as RNN and LSTM can expand the forward direction in the forward and backward directions and both directions through a modified algorithm called bidirectional LSTM to achieve additional performance shapes. However, the proposed digital unit can also be expanded in the reverse and both directions in the same way as the bidirectional. The proposed digital unit is different from RNN and LSTM in that composite input features are combined. Video scene frames are clustered in characteristic patterns and may be clustered in the forms of operation A, operation B, and operation C. Surgical operations and / or special operations of specific characters in the game may all be re-associated by a series of sequential pattern of learned operation clusters (A, B, C,...). This means that in one embodiment of the present invention, when sequential related patterns such as A→B→D and A→B→F are frequently observed with high sequential relevance in the training data, the pattern will be learned along with the sequence. Specific repetitive operations may be learned and reproduced by the sequential related patterns of video clusters. The reproduction described here means that when a partial pattern of the above-mentioned part is input by the input, it is possible to analogically infer by the relevant rules what kind of operation pattern of the cluster the subsequent pattern is.
[0401] Various embodiments explain the relevant rules in the reverse time series direction.
[0402] Using the first type of GNN regression model, the output (Output) data (traces and results), which is still image information, is predicted. The predicted value of the second type 1205 of the first GNN regression model, where the vector product of the point cloud that causes the output data (traces and results) is in the reverse direction. Analyze the result value to find the vector product of the point cloud on the GNN framework due to the change in time and return it to the platform user. Discover the rule that if there is a result value, there is a vector product of the point cloud. When the first type of GNN regression model presents any predicted value (result value) and / or still image information, the second type of GNN regression model returns the vector product (cause value) of the point cloud due to the change in time. To search for the meaningful relationship between result values for related rule inference, construct a set of data sets of result values and a set of transactions that return the vector product (cause value) of the point cloud. Related rules have antecedent events and consequent events, which are respectively included in the sets of result values and cause values. This is obtained as a result of related rule inference. Since the vector product is information with complexity, there are many related rules. An evaluation criterion is needed to find meaningful related rules. As evaluation metrics, support, confidence, lift, etc. are used. In the related rule algorithm, each set of result values and cause values means a cluster of still image information and a cluster of video information in digital units 4 and 5, respectively.
[0403] The time series segmentation selection labeling function described above will be further explained.
[0404] Without hierarchical labeling, the server 200 can return to the user the segmentation time points of the first video information and the second video information (for example, including still image information, label values of still image information, etc.), and the user can perform time series segmentation selection labeling or body part-specific selection labeling on the returned value (segmentation time point). When time series segmentation selection labeling or body part-specific labeling is performed, digital unit 3 or digital unit 4 or digital unit 5 is generated.
[0405] The video information is the first video information 1503 or the second video information 1506, and the step of receiving the second selection labeling 1208 or the third selection labeling information includes the time-series division selection labeling 1701 in FIG. 17. The video information is the repeated predicted value of the GAN and / or GNN prediction model 1605, that is, "the first, second, third,... video information".
[0406] The second hierarchical labeling 1220 or the third hierarchical labeling information including the time-series division selection labeling is related to the hierarchical clusters in FIG. 8 or FIG. 10.
[0407] In an embodiment of the present invention, even when the hierarchical clustering labeling is omitted, the time-series division selection labeling 1701 may be included in and executed by the second selection labeling 1208 or the third selection labeling, or may be executed before and after the second selection labeling 1208 or the third selection labeling.
[0408] The second hierarchical cluster 1507 or the third hierarchical cluster is processed (or computerized) based on the digital unit 3 or the digital unit 4 in FIG. 8 or FIG. 10.
[0409] FIG. 8 shows a hierarchical cluster 800 processed based on the digital unit 3, and FIG. 10 shows a hierarchical cluster (1000) processed based on the digital unit 4.
[0410] For "the predicted values of the first and second GNN regression models of type 1 or the predicted values of the first and second related rules of type 1" or "the predicted values of the first and second GNN regression models of type 2 or the predicted values of the first and second related rules of type 2", the method in which the user attaches the labels of ACCEPT or REJECT to select or reject the division time point (still video information or the label value of the attribute) is defined as the "time-series division selection labeling 1701" in FIG. 17.
[0411] After the user performs the time-series segmentation selection labeling 1701 in FIG. 17 with reference to the label classification, the second classification model 1209 is induced and / or inferred.
[0412] When the artificial intelligence learns the segmentation time point and returns it to the user, the user makes a selection through the ACCEPT button or the REJECT button. The classification model classifies the labeled information again, and the "GNN regression model or related rules (including deep learning)" returns the collective-intelligence prediction value. The user repeatedly performs the time-series segmentation selection labeling 1701 in FIG. 17. The user who refers to the label classification inputs the label value through the execution result screen (or view screen) of the application of the terminal 100.
[0413] In one embodiment of the present invention, when the user presses the REJECT button in the time-series segmentation selection labeling 1701, a table (or cursor) indicating the time is moved on the timeline within the playback bar on the execution result screen (or view screen) of the application of the terminal 100, and after capturing the still image information at the point in time when the video is to be split, the label value labeled by time-series segmentation is directly input by pressing the ACCEPT button.
[0414] The digital unit 3 1705 and the digital unit 4 1706 used in the induction and / or inference algorithm 1200 in FIG. 12 are the divided rectangular parallelepipeds 301 in FIG. 3.
[0415] The attribute in the digital unit 3 1705 is the still image information 403 at the end of the k-th step of the divided motion video such as an avatar, a human, or a robot, which is the black rectangle in FIG. 4.
[0416] Referring to FIG. 4, the still image information at the end of the n-th step of the divided motion video also corresponds to the attribute, which is the last black rectangle in FIG. 4.
[0417] The attribute in digital unit 4 1706 is the still image information 503 at the end of the L-th step of the segmented motion video such as an avatar, a human, or a robot, which is the black rectangle in FIG. 5.
[0418] Referring to FIGS. 4 to 5, the still image information at the end of the (k, L)-th step of the segmented motion video also corresponds to the attribute, which is the last black rectangle in FIGS. 4 to 5.
[0419] FIGS. 8 and 10 are K3 and K5 clusters based on digital unit 3 1705 and digital unit 4 1706.
[0420] In various embodiments, the still image information at the start portion is also an attribute, and it becomes digital unit 3 or digital unit 4 by the sum with the video information that is the target attribute.
[0421] In one embodiment of the present invention, FIGS. 8 or 10 is a hierarchical clustering pedigree diagram created by the label value attached when the steps of the video are segmented according to the variable value input into the input window of the execution result screen (or view screen) of the application of FIGS. 7 to the terminal 100.
[0422] The time-series segmentation selection labeling 1701 in FIG. 17 segments the video based on "data unit 3 1703 or data unit 4 1704" and "digital unit 3 1705 or digital unit 4 1706".
[0423] In one embodiment of the present invention, digital unit 3 1705 or digital unit 4 1706 is obtained by segmenting the video in parallel with the time-series segmentation selection labeling 1701 and data alignment. The user attaches an ACCEPT label or a REJECT label to the video or the label value indicating the order of the video to select or reject the order of the segmented videos in order to align the order of the aligned videos.
[0424] The digital unit 4 is motion video information segmented by the time-series segmentation selection labeling 1701 in which the user refers to the label classification and segments the motions of an avatar, a human, a robot, etc. into characteristic detailed motions around 0.5 seconds to 3 seconds. The digital unit 5 enables further segmentation of the video compared to the digital unit 4.
[0425] The digital unit 3 is motion video information segmented by the time-series segmentation selection labeling 1701 in which the user refers to the label classification and segments the motions of an avatar, a human, a robot, etc. into characteristic motions around 3 seconds to several tens of seconds.
[0426] The data unit described in the embodiments of the present invention is a unit of a composite feature vector generated by the user, and the digital unit is a unit of a composite feature vector generated by the interaction between the user and artificial intelligence.
[0427] In one embodiment of the present invention, the label classification that classifies the motions of an avatar, a human, a robot, etc. into characteristic detailed motions around 0.5 seconds to 3 seconds is the above [Table 5] or [Table 10].
[0428] In one embodiment of the present invention, the data unit 3 and the digital unit 3 may be units of video information segmented into units from several seconds to several tens of seconds. When a quantum cloud computing device with 3000 qubits or more is commercialized and the computing power is improved far more than it is now, the data unit 3 and the digital unit 3 are used for the generation and output of video information.
[0429] The digital unit 3 1705 is obtained by processing the sum of the attribute (still video information) and the target attribute (video information) in the same manner as the data unit 3 1703.
[0430] The digital unit 4 1906 is obtained by processing the sum of the attribute (still video information) and the target attribute (video information) in the same manner as the data unit 4 1904.
[0431] In one embodiment of the present invention, a large number of users (such as students at the Air Force Academy and / or fighter pilots) imitate and operate for about one minute the scene of flying a fighter plane in the movie "Top Gun" using a virtual fighter simulator (VEHICLE VR simulator), proceed with labeling, and obtain data sets for each data unit and digital unit to be used in the guidance and / or inference algorithm of FIG. 12. Since there are characteristic operations for each method in the detailed flight maneuvers of flying a fighter plane, if a large number of users perform similar virtual flights, the entire video is divided into short videos about 1 to 2 seconds long.
[0432] In one embodiment of the present invention, a large number of users use a VR treadmill to fire an electric controller-type weapon for about one minute of the battle scene in the movie "Saving Private Ryan" and proceed with labeling. The movements of infantry and engineers in the movie (continuous actions such as firing a handgun and throwing a grenade) can also be divided into short videos about 1 to 2 seconds long.
[0433] In one embodiment of the present invention, the time series segmentation method for digital unit 31905 is as follows.
[0434] Referring to label classifications such as [Table 1] to [Table 4] above, perform labeling for dentists, doctors, etc. When the guidance and / or inference algorithm 1200 in FIG. 12 is advanced, the "GNN regression model" returns still image information and returns the segmentation time point and label values (s1, s2, s3, k) to platform users (doctors, dentists, etc.). For the returned values, the user performs time series segmentation selection labeling 1701.
[0435] In one embodiment of the present invention, the following [Table 12] explains the 30-second video of proceeding with the deletion of laminate No. 11 (tooth formula) of the maxillary central incisor, which is divided into 10 steps and divided at intervals of about 2 to 4 seconds for 30 seconds. The video can be segmented in digital unit 4 by the user's time series segmentation selection labeling 1701.
[0436]
Table 12
[0437] Even if the video information for a number of patients (avatars and digital cards) divided into 10 steps as described above belongs to the same specific cluster in FIG. 9, the detailed surgical procedures and the order of the procedures in the corresponding video may vary depending on the medical skills of the operating doctor. Other orders may be pre-processed based on the label order in [Table 5] above and applied to the classification model. Also, for the steps with different orders, omitted parts, and / or added parts, the video information is aligned and clustered based on the label order in [Table 12].
[0438] Referring to the label classification ([Table 12]), when the dentist performs the labeling, for the feedback of the artificial intelligence, the dentist attaches an ACCEPT label or a REJECT label. When performing time-series segmentation selection labeling for the segmentation point (still image information), the classification model re-classifies the labeled information, and the "GNN regression model" further returns the segmentation point (still image information) and the label value that have been made more collective. For the feedback of the artificial intelligence, the dentist selects by pressing the ACCEPT button or the REJECT button. When the dentist performs the labeling in the above manner, while the GNN regression model that divides the video information and returns the still image information returns the still image, the guidance and / or inference algorithm 1200 in FIG. 12 returns the segmentation point (attribute value) and the label value to the dentist. For the predicted value of the artificial intelligence, if the dentist presses the ACCEPT button or the REJECT button to attach a label and selects or rejects the segmentation point (attribute or the label value of the attribute), the classification model re-classifies the labeled information, and the "Type 1 GNN regression model" further returns the segmentation point (still image information) and the label value that have been made more collective. Eventually, a sufficiently collective digital unit 4 1906 is generated.
[0439] The selection labeling function for each body part described above will be further explained.
[0440] The video information is the first video information 1503 or the second video information 1506, and the step of receiving the second selection labeling 1208 and the third selection labeling information includes the selection labeling 1702 for each body part.
[0441] The video information is the repeated predicted values of the GAN and / or GNN prediction model 1605, which are "the first, second, third,... video information".
[0442] Performing the selection labeling for each body part, the second hierarchical labeling 1220 or the third hierarchical labeling information including the selection labeling information for each body part becomes a hierarchical cluster.
[0443] In an embodiment of the present invention, even when the hierarchical clustering labeling is omitted, the selection labeling 1702 for each body part may be included and executed in the second selection labeling 1208, or may be executed before and after the second selection labeling 1208.
[0444] The second hierarchical cluster 1507 or the third hierarchical cluster is a cluster computerized based on the digital unit 5.
[0445] The digital unit 5 1707 is processed by adding the attribute (still video information) and the target attribute (video information) in the same manner as the digital unit 4 1706.
[0446] The data unit 3 1703, the data unit 4 1704, the digital unit 3 1705, or the digital unit 4 1706 is processed by the selection labeling 1702 for each body part as "the digital unit 5 1707".
[0447] In order to determine the operation order for each body part from the operations of an avatar, a human, a robot, etc., labels for determining the order are attached to each body part, and labeling is performed to change the operation order in the actual video, which is defined as "selection for each body part".
[0448] In an embodiment of the present invention, after performing "selection by body part", it is possible to select or reject by attaching an ACCEPT label or a REJECT label to preprocessing operations (such as deletion and addition) by arranging video data.
[0449] After the user refers to the label classification and performs selective labeling 1702 by body part on the first and second videos (or video information), the second and third classification models are induced and / or inferred.
[0450] A step in which the user refers to the label classification and performs selective labeling 1702 by body part on the first video information 1503 or the second video information 1506, whereby the video is divided in digital units 5 1707.
[0451] The attribute in the digital unit 5 1707 is still image information at the end of the f-th step of the divided motion video such as an avatar, a human, or a robot, and is the black rectangle in FIG. 6.
[0452] Referring to FIG. 6, the still image information at the end of the f-th step of the divided motion video also corresponds to the attribute and is the last black rectangle in FIG. 6.
[0453] FIG. 11 shows K6 clusters based on the digital unit 5 1707.
[0454] The digital unit 5 1707 used in the induction and / or inference algorithm 1200 of FIG. 12 is the divided rectangular parallelepiped 301 of FIG. 3.
[0455] In various embodiments, the still image information at the start portion is also an attribute, and together with the video information that is the target attribute, it becomes the digital unit 5.
[0456] In one embodiment of the present invention, FIG. 11 is a phylogenetic diagram of hierarchical clustering created by label values attached when video steps are divided by variable values (label values) input into the input window of the execution result screen (or view screen) of the application of the terminal 100.
[0457] The motion videos of avatars, humans, robots, etc. are divided in digital unit 5 by "Selective Labeling 1702 by Body Part".
[0458] In one embodiment of the present invention, regarding the removal of the 11th laminate tooth, most dentists create an index for tooth removal and remove the tooth, but there are also dentists who do not create an index, and there are also people who do not use a depth gage bur. Hierarchical clustering is performed based on the above differences to align and preprocess video information. There may be a person who does not use the order when removing teeth (such as the order of the tooth root, the center, and the cutting part) as a criterion for label classification and proceeds in his own order. In such a case, a label classification that matches the labeling order is created by specifying the order of the videos divided by the labeling that specifies the detailed order for body parts such as the center, cutting part, and tooth root of the upper central incisor. Also, the video information is aligned in the order of the above labeling. Although included in the same specific cluster (in one embodiment of the present invention, the method of removing the 11th tooth without using an index), the order of removing the teeth (the order of removing the tooth root, the center, and the cutting part) is different from other videos and still image information. By preprocessing operations such as the order labeling of body parts and the alignment of video information, the error value of the classification model is reduced and the accuracy of the classification model is increased.
[0459] In one embodiment of the present invention, when a dentist removes teeth in the order of the gingival part, the central part, and the cutting part, and when a dentist removes teeth in the order of the central part, the cutting part, and the gingival part, the video information is aligned in a manner such as the deletion order of the gingival part, the central part, and the cutting part, and is segmented and clustered in this order. Further, if the dentist uses the mouse cursor to point to a specific part of the body or even considers a specific part of the body, the artificial intelligence returns the boundary lines and boundary surfaces of the part pointed to or considered by the user through object recognition. Further, the artificial intelligence returns the aligned information regarding the treatment order to the user. In contrast, the user determines whether "the part intended or considered by oneself is correct or incorrect" and / or "the order intended by oneself is correct or incorrect" and / or "the label value for the order is correct or incorrect". Only by making a judgment using the brain-computer interface in this way, the video and still images are labeled with an ACCEPT label or a REJECT label and aligned. The above labeling is repeated and applied to the induction and / or inference algorithm 1200 of FIG. 12.
[0460] In one embodiment of the present invention, as shown in [Table 6], there are about 28 teeth in the oral cavity of a healthy adult, and each tooth has a dental formula (tooth number). The upper right central incisor is No. 11. When performing a surgical procedure to remove four teeth with dental formulas 22, 21, 11, and 12 for laminate treatment, when all dentists remove teeth for laminate treatment, they do not proceed in a fixed order of tooth numbers (dental formulas). Therefore, the above video information is aligned in a fixed order (dental formula) and pre-processed for the deleted or added video information.
[0461] When performing time-series segmentation labeling, if hierarchical clustering and alignment (dental formula order) for the specific treatment order of body parts are carried out simultaneously, more accurate clustering is possible (hierarchical clustering and selection by body part). Further, when attempting to perform more detailed selection labeling 1702 by body part, the method and order of removing the laminate teeth of the upper central incisor (tooth No. 11) can also vary among dentists. Therefore, the video information is aligned based on label classification (certain criteria) and pre-processed for the deleted or added video information.
[0462] In one embodiment of the present invention, when attempting to obtain a video divided into digital units with an instantaneous small data size of 0.5 seconds or less in a meta bus soccer game video, the video is subdivided and divided using the selection labeling 1702 by body part and alignment. When Son Hoon Min performs a long in-step dribble of 3 steps, the soccer ball instantaneously touches Son Hoon Min's foot in the order of toe touch of the preset label classification, 1-step race, ankle touch, and 2-step race. However, if a specific user who reproduced this touches and races in the order of ankle touch, 1-step race, toe touch, and 2-step race, for the touch order and race order of the specific user's in-front dribble, a selection labeling 1702 by body part is performed using a brain-computer interface.
[0463] In one embodiment of the present invention, in [Table 10], if the k-th movement of Jennie in the opened concert broadcast on July 8, 2022 is a front and back wave, in [Table 11], Jennie of Blackpink's front and back wave movement raises or moves back and forth in the order of raising the left arm, raising the right arm, moving the chest, moving the abdomen, moving the pelvis, and moving the feet. If a specific user makes a front and back wave in the order of moving the feet, moving the pelvis, moving the abdomen, moving the chest, raising the right arm, and raising the left arm, a selection labeling (1902) by body part is performed using a brain-computer interface. The movement video of the specific user is aligned in the order of Jennie of Blackpink's movements. Also, the video of 3 minutes and 14 seconds in total can be time-series divided into about 200 videos of around 1 second to 2 seconds. The dance movements are continuous combinations of movements of the head, hands, feet, and torso. Instead of the selection labeling 1702 by body part, a time-series division selection labeling 1701 may be performed.
[0464] In addition, the server 200 repeatedly performs the selection labeling process, classification model inference process, prediction model inference process, additional selection labeling process for the generated first video, additional classification model inference process, and additional prediction model inference process on the plurality of raw data provided from the plurality of terminals 100 in relation to the corresponding specific topic, and generates (or updates) a second video with collective wisdom in relation to the corresponding specific topic (or in relation to a comparison target video related to the corresponding specific topic).
[0465] At this time, the server 200 can also provide the second video that was last updated (or newly generated) to the plurality of terminals 100 that provided the raw data in relation to the corresponding specific topic in real time or in response to a request from a specific terminal 100.
[0466] As a result, all the terminals 100 or a specific terminal 100 that provided the raw data related to the corresponding specific topic to the server 200 can be provided with the latest second video with collective wisdom related to the corresponding specific topic.
[0467] The "video information of the first, second, third,..." repeatedly generated by the GAN and / or GNN prediction model 1605 is repeatedly learned with the basic video information 1601 in a single model. Hierarchical labeling and selection labeling 1604 are repeatedly executed. The classification model is repeatedly inferred, and the GAN and / or GNN prediction model 1605 is repeatedly induced and / or inferred.
[0468] In one embodiment of the present invention, time series segmentation selection labeling 1701 and / or body part-specific selection labeling 1702 are repeatedly executed.
[0469] In addition, the server 200, in conjunction with the terminal 100, collects operation-related videos (or operation-related videos related to at least one of a human, an avatar, and an item), such as actual humans (or actual people) output from (or managed by) the terminal 100, virtual avatars, and items, and meta information related to the corresponding operation-related videos, in relation to a specific topic. Here, the specific topic (or specific content) includes medical practices (including, for example, treatments, surgeries, etc.), dance, sports events (including, for example, soccer, basketball, table tennis, etc.), games, e-sports, and the like. Further, the operation-related video (or basic video information / raw data) related to the human may be a video obtained (or captured) of the actions (or operations / acts) performed by an actual human (or person / influencer) in relation to the specific topic. Also, the operation-related video of the avatar and / or item may be a video generated by a selection labeling process, a classification model inference process, a prediction model inference process, etc. based on any raw data related to the corresponding specific topic.
[0470] In one embodiment of the present invention, the visual data 1801 of the avatar, item, and human operation in FIG. 18 includes time data of the movement of a vehicle operated by the avatar or the human. Here, the visual data 1801 indicates raw data for the actions of a user (or human) in the real world.
[0471] In addition, in order to embody the collected operation-related video as the actual operation of the robot, the server 200 reconstructs the collected operation-related video (or the operation-related video of the actual human, virtual avatar, item, etc. collected) as a robot operation video. Here, the robot is a robot arm manufactured in a form that can operate in a tooth extraction VR simulator using the visual data of the tooth extraction VR simulator, a robot arm manufactured in a form that can operate in a surgical VR simulator using the visual data of the surgical VR simulator, a robot manufactured in a VEHICLE form using the visual data of a VEHICLE VR simulator, and a humanoid robot that can operate on a VR treadmill.
[0472] That is, the server 200 converts the coordinate information related to the actual human, virtual avatar, item, etc. included in the corresponding operation-related video into robot coordinate information for applying the operation of the corresponding actual human, virtual avatar, item, etc. to the actual robot based on the collected operation-related video, the meta information related to the operation-related video, etc., and reconstructs the corresponding operation-related video as the robot operation video.
[0473] In addition, the server 200 transmits the robot operation video (or the reconstructed robot operation video), the meta information related to the corresponding robot operation video, the collected operation-related video, the meta information related to the operation-related video, the comparison target video retrieved in relation to the collected operation-related video (or the robot operation video) among the plurality of comparison target videos managed by the server 200, the meta information related to the corresponding comparison target video, etc. to a selected specific terminal 100 among the plurality of terminals 100 pre-registered in the server 200.
[0474] Further, the specific terminal 100 receives the robot operation video transmitted from the server 200, meta information regarding the corresponding robot operation video, the operation-related video, meta information related to the operation-related video, a comparison target video corresponding to the operation-related video (or robot operation video), meta information related to the corresponding comparison target video, and the like.
[0475] While accurately measuring the spatial and temporal coordinates of the movement of the robot in connection with the terminal 100, the movement of the robot is evaluated via the display device 1808 and the user interface 1809 of the execution result screen (or view screen) of the application of the terminal 100, enhanced in a manner in which the user performs robotics selection labeling 1810, and defined as "collective intelligence robotics 1803" which operates 1806 on the server 200 and the robotics programming inferred by the collective intelligence model. For the basic robotics video information, selection labeling (basic selection labeling) is performed before the first robotics selection labeling to infer and / or induce the first collective intelligence robotics 1803. For the basic robotics video information, hierarchical labeling and / or selection labeling may be performed in the same manner as the method of FIG. 12.
[0476] In an embodiment of the present invention, the visual data output from the execution result screen (or view screen) of the application of the terminal 100 is an operation screen in virtual reality, augmented reality, mixed reality, extended reality, etc. of the robot provided by the terminal 100.
[0477] The robotics video information 1813 corresponds to the attributes and target attributes of FIGS. 3 to 6 above and is visual data generated by the collective intelligence robotics 1803.
[0478] The first collective intelligence robotics 1902 is programmed with the input of the first basic robotics video information 1901 to generate the first robotics video information 1911.
[0479] The basic robotics video information 1802 in FIG. 18 is the robot motion data (video information) reconstructed by the server 200 from the motion data 1801 of an avatar, human, robot, etc. secured from the terminal 100 in a synchronized state such that the motion information and position information of the metabus user are aligned with the coordinates of the virtual environment. The visual data 1801 of the avatar motion secured from the terminal 100 is the predicted value of the GAN and / or GNN prediction model 1605 in FIG. 16, the "first, second, third,... video information" in FIG. 15, and hereinafter, the predicted value of the GAN and / or GNN prediction model 1605 of the metabus world that is repeated. Alternatively, the visual data 1801 of the human (or user) motion means the raw data for the user's motion in the real world. The visual data 1801 of the human (or user) motion is reconstructed by the server 200 as the robot motion data (video information), and the robot motion video information reconstructed from the visual data of the human motion is included in the basic robotics video information 1802.
[0480] In one embodiment of the present invention, the visual data output from the execution result screen (or view screen) of the application of the terminal 100 is the motion screen of the robot 1807 provided by the terminal 100, which may be virtual reality, augmented reality, extended reality, mixed reality, etc.
[0481] In order to reduce the error between the coordinate system on the user interface 1809 of the terminal 100 and the coordinate system in the robot motion, an actual distance coordinate system based on the size of the robot is estimated, and the angles of each robot joint are extracted and controlled.
[0482] In an embodiment of the present invention, the visual data of a tooth removal VR simulator, a surgical VR simulator, a VEHICLE VR simulator, and a VR treadmill are used to manufacture a robot in the form of a robotic arm, a humanoid, or a VEHICLE.
[0483] Referring to FIG. 18, the basic robotics video information 1802 is input into the collective intelligence robotics 1803. The GAN and / or GNN robotics prediction model included in the collective intelligence robotics 1803 includes an interface API process, and the prediction model outputs the robotics video information 1813 to the execution result screen (or view screen) of the application of the terminal 100. The GAN and / or GNN robotics prediction model is a model of visual data related to the robotics operation in the same manner as the GAN and / or GNN prediction model 1605. The robotics video information 1813 is the "first, second, third,... robotics video information" repeatedly output and / or generated by the GAN and / or GNN robotics prediction model. Repeated robotics selection labeling 1810 is performed on the video information.
[0484] The visual data output from the collective intelligence robotics 1803 is transmitted to the robot simulation engine 1804, and the robot is operated 1806 via the API communication 1805 with the robot, and is output via the display device 1808 and the user interface 1809 through the graphics engine 1807.
[0485] In one embodiment of the present invention, the programming of the robotics in the server 200 is as follows. The vision sensor is interfaced with ROS using ROS (Robot Operating System), OpenCV (Open Source Computer Vision), and PCL (Point Cloud Library), and programming is performed using libraries such as OpenCV and PCL.
[0486] In one embodiment of the present invention, in the metabus hospital and dental hospital game, the terminal 100 creates a 3D model related to the patient's medical history and aligns the 3D patient coordinate system based on the position and state of the medical history and the video information with the coordinate system of the patient placed on the operating table.
[0487] Thus, according to the present invention, in a hospital game, a dental user can be provided with a service of substituting items such as medical devices, medical equipment, and materials into the face and body of a digital avatar (patient's avatar) and generating and / or outputting them in various combinations.
[0488] In one embodiment of the present invention, when sufficient visual data regarding virtual surgery and tooth extraction procedures is secured via a VR simulator operated by a doctor and / or a dentist, it is possible to fabricate an artificial intelligence surgical and treatment robot capable of performing automated surgery and treatment with the VR simulator through robotics programming. In the utilization of clusters collected from "actual medical institution data collection and virtual dental simulator and virtual surgery simulator", an initial model of "artificial intelligence capable of performing automated surgery and treatment" is created through a sequential model in which related rules and those predicted previously are repeated as the next input. The artificial intelligence is enhanced by robotics selection labeling 1810. The artificial intelligence advances the virtual surgery and treatment, and in response, the doctor advances the labeling and applies the guidance and / or inference algorithm 1200 of FIG. 12 to enhance the artificial intelligence.
[0489] The first robotics video information 1911 in FIG. 19 is the second attribute and the second target attribute in FIGS. 4 to 6. The basic robotics video information 1802 is data belonging to a specific cluster, which is one of the clusters in FIGS. 7 to 11. The "robotics video information 1813" also becomes data belonging to the same specific cluster.
[0490] In addition, the server 200 performs selective labeling on the robot operation video. Here, the selective labeling (or selection labeling) refers to a labeling method of setting (or attaching) a label (or label value) for the presence or absence of an error (or abnormality) at a specific time point (or specific section) of the robot operation video. At this time, for the time point (or section) of the robot operation video for which no label (or label value) has been set by the selective labeling, a preset default label value (for example, an approval label) may be set.
[0491] That is, the server 200, in conjunction with the terminal 100, responds to the input of the user of the corresponding terminal 100 (or the user's selection / touch / control) for the robot operation video displayed on the corresponding terminal 100, and sets (or receives / inputs) the label (or label value) at a specific time point (or specific section) of the corresponding robot operation video.
[0492] In addition, the server 200 receives one or more selective label values at one or more specific time points (or specific sections) related to the robot operation video transmitted from the terminal 100, the meta information of the corresponding robot operation video, the identification information of the corresponding terminal 100, and the like.
[0493] In the embodiments of the present invention, mainly described is that, at the terminal 100, in response to the input of the user, one or more selective label values are set (or received / input) at one or more specific time points (or specific sections) of the corresponding robot operation video. However, the present invention is not limited thereto. The server 200 performs a video analysis function on the corresponding robot operation video and a comparison target video related to the corresponding robot operation video, and based on the execution result of the video analysis function, automatically sets one or more selective label values at one or more specific time points (or specific sections) for the corresponding robot operation video, respectively.
[0494] Also, in the server 200, when one or more selection label values are set for the corresponding robot operation video at one or more specific time points (or specific intervals), the server 200 provides information regarding the one or more selection label values at the one or more specific time points (or specific intervals) related to the set corresponding robot operation video to the terminal 100. In the corresponding terminal 100, information regarding the one or more selection label values at the one or more specific time points (or specific intervals) related to the corresponding robot operation video set by the server 200 is displayed, and in response to the input of the user of the corresponding terminal 100, it may be configured to determine the presence or absence of final approval for the one or more selection label values at the corresponding one or more specific time points (or specific intervals).
[0495] At this time, before or after performing selection labeling on the corresponding robot operation video, the server 200 performs hierarchical labeling on the corresponding robot operation video in conjunction with the terminal 100. Before / after performing hierarchical labeling, it is also possible to perform selection labeling on the corresponding robot operation video. Here, the hierarchical labeling (or hierarchical leveling) is input feature engineering by the user, which indicates a labeling method of attaching labels indicating features of the corresponding robot operation video and dividing (or classifying) the corresponding robot operation video into a plurality of sub-robot operation videos according to the features.
[0496] That is, the server 200, in conjunction with the terminal 100, with respect to the robot operation video displayed on the corresponding terminal 100, in relation to the corresponding specific topic, refers to (or based on) a plurality of preset label classifications, and in response to the input (or user selection / touch / control) of the user of the corresponding terminal 100, sets (or receives / inputs) labels (or label values) at other specific time points (or other specific intervals) of the corresponding robot operation video.
[0497] Also, the server 200 divides the robot operation video into a plurality of sub-robot operation videos.
[0498] In the embodiments of the present invention, in the terminal 100, mainly described is setting (or receiving / inputting) one or more hierarchical label values at one or more other specific time points (or other specific intervals) among the corresponding robot operation videos according to user input. However, it is not limited thereto. The server 200 performs a video analysis function on the corresponding robot operation video and a comparison target video related to the corresponding robot operation video, and based on the execution result of the video analysis function, may automatically set one or more hierarchical label values for the corresponding robot operation video at one or more other specific time points (or other specific intervals) respectively.
[0499] Also, in the server 200, when one or more hierarchical label values are set for the corresponding robot operation video at one or more other specific time points (or other specific intervals), the server 200 provides information on the one or more hierarchical label values at one or more other specific time points (or other specific intervals) related to the set corresponding robot operation video to the terminal 100. In the corresponding terminal 100, information on the one or more hierarchical label values at one or more other specific time points (or other specific intervals) related to the corresponding robot operation video set by the server 200 is displayed, and it may be configured to determine the presence or absence of final approval for the one or more hierarchical label values at the corresponding one or more other specific time points (or other specific intervals) according to the input of the user of the corresponding terminal 100.
[0500] Receive first robotics selection labeling 1903 information for the first robotics video information 1911 which is the video information.
[0501] The induction and / or inference method of the first robotics classification model 1904 is the same as the method of the second classification model in FIG. 12.
[0502] Perform the first robotics selection labeling 1903 on the first robotics video information 1911 which is the output of the first collective intelligence robotics 1902. Define the classification model obtained by classifying the visual data acquired by performing the first robotics selection labeling 1903 as the "first robotics classification model 1904". Hereinafter, the robotics classification model 1911 is repeated. The robotics selection labeling 1910 is in the same manner as the selection labeling 1604 in FIG. 16.
[0503] For the operation of the robot output from the user interface 1809 in FIG. 18 in the form of the execution result screen (or view screen) of the application of the terminal 100, the user performs the robotics selection labeling 1810 in FIG. 18. The operation of the robot is the robotics video information 1813.
[0504] The robots of the initial model of the "collective intelligence robotics 1803" may have many errors in operation. For the somewhat inaccurate movements of the robot 1809, the robotics developer performs supervised learning by means of the robotics selection labeling 1810 and classification. Provide the avatar, items, and spatial environment, narration, etc. generated and output in the virtual simulation to the collective intelligence robotics 1803, and perform supervised learning on the artificial intelligence by the method in which the user performs the robotics selection labeling 1810. The guidance and / or inference algorithm 1200 in FIG. 12 enhances the collective intelligence robotics 1803.
[0505] In addition, the server 200 performs machine learning on the artificial intelligence infrastructure based on information about the selected-labeled robot operation video, etc., and generates (or confirms) a classification value for the corresponding robot operation video based on the result of the machine learning. Here, the classification value for the corresponding robot operation video (or the classification value of the corresponding robot operation video / the classification value of the selected-labeled robot operation video / the classification value of the hierarchically labeled robot operation video) may be a value obtained by classifying selection labeling values, hierarchical labeling values, etc. by the same item.
[0506] That is, the server 200 performs machine learning (or artificial intelligence / deep learning) using information about the selected and labeled robot operation video as input values of a preset classification model, and generates (or confirms) a classification value for the corresponding robot operation video based on the result of the machine learning (or the result of artificial intelligence / the result of deep learning).
[0507] In FIG. 18, the video information output to the user interface 1809 is labeled by the robotics selection labeling 1810 and classified by the robotics classification model 1811. The classified visual data is the labeled robotics label information 1812 transmitted to the collective intelligence robotics 1803.
[0508] In an embodiment of the present invention, the robotics selection labeling 1810 includes hierarchical labeling in the same manner as video processing of the metabus, time-series segmentation selection labeling 1701, selection labeling by body part 1702, and the like.
[0509] The information classified by the first robotics classification model 1904 is the first robotics label information 1905.
[0510] In one embodiment of the present invention, a number of users corresponding to experts in each field of virtual simulation games proceed with robotics selection labeling 1810 via the interface 1809 of the terminal 100. When sufficient visual data is secured, an artificial intelligence robot that operates a VR simulator is manufactured. When an initial model of the collective intelligence robotics 1803 that operates the VR simulator using robot joints, arms, legs, etc. is developed, the capabilities of the collective intelligence robotics 1803 model are enhanced by supervised learning based on the labeling by the users. When enhanced, an initial model of the collective intelligence robotics 1803 that can operate in the actual real world can be developed. Also in this case, the collective intelligence robotics 1803 is enhanced by supervised learning based on the robotics selection labeling 1810 of the users and experts. The collective intelligence robotics 1803 in FIGS. 18 and 19 repeatedly applies the induction and / or inference algorithm 1200 in FIG. 12 to repeat the labeling and enhance the model of the collective intelligence robotics 1803. By enhancing the collective intelligence robotics 1803, an initial model of artificial intelligence that can perform automated surgery and operations in an actual medical field using a robot arm can be developed. Also in this case, the automation capabilities of the initial model of artificial intelligence are enhanced by supervised learning based on the robotics selection labeling 1810 of the doctors.
[0511] In one embodiment of the present invention, the collective intelligence algorithm evaluated and trained by a user (doctor) enhances artificial intelligence inference to a level where virtual surgery simulation and tooth extraction simulation can be automated without errors and artificial intelligence is enhanced to the level of automated surgery. When an initial model of artificial intelligence capable of automated surgery and operation on a VR simulator using a robotic arm is developed, the automation ability of the artificial intelligence model is enhanced by supervised learning through the doctor's robotics selection labeling 1810. When enhanced, an initial model of artificial intelligence capable of automated surgery and operation in an actual medical field using a robotic arm can be developed. In this case as well, the automation ability of the initial model of artificial intelligence is enhanced by supervised learning through the doctor's robotics selection labeling 1810.
[0512] In one embodiment of the present invention, a humanoid robot using a robot head, a robot arm, robot legs, a robot body, robot joints, etc. is manufactured, and vehicle robots such as autonomous vehicles, drones, airplanes, etc., an artificial intelligence dentist robot, and an artificial intelligence doctor robot are manufactured.
[0513] Also, the server 200 uses the classification value for the generated corresponding robot operation video (or the classification value of the corresponding robot operation video), information regarding the selected and labeled robot operation video, the corresponding robot operation video, meta information related to the corresponding robot operation video, the comparison target video, meta information related to the corresponding comparison target video, etc. as input values to perform machine learning (or artificial intelligence / deep learning), and based on the result of machine learning (or the result of artificial intelligence / the result of deep learning), a first robotics video corresponding to the corresponding robot operation video is generated. At this time, the first robotics video may be an operation-related video of an avatar, an item, a robot, etc. generated based on the robot operation video, a video in which the robot operation video is updated, or the like.
[0514] That is, the server 200 performs machine learning (or artificial intelligence / deep learning) using, as input values for a preset prediction model, the classification value for the generated corresponding robot operation video (or the classification value of the corresponding robot operation video), information regarding the selected and labeled robot operation video, the corresponding robot operation video, meta information associated with the corresponding robot operation video, the comparison target video, meta information associated with the corresponding comparison target video, etc., and generates a first robotics video associated with the corresponding robot operation video based on the result of the machine learning (or the result of the artificial intelligence / the result of the deep learning).
[0515] Further, the server 200 transmits (or provides) the generated first robotics video to the terminal 100.
[0516] The second robotics video information 1912 is a video that is a predicted value of the second collective intelligence robotics 1906, and is defined as a predicted value of a prediction model advanced by repeated application of the repeated induction and / or inference algorithm 1200 in the present invention.
[0517] The second robotics video information 1912 is the third attribute and the third target attribute in FIGS. 4 to 6.
[0518] The first robotics label information 1905 classified by the first robotics classification model 1904 is input into the second collective intelligence robotics 1906, and the second basic robotics video information 1907 is also input into the second collective intelligence robotics 1906 and programmed as a single model to generate the second robotics video information 1912.
[0519] Further, the server 200 performs additional selection labeling on the first robotics video. Here, the additional selection labeling (or additional selection labeling) refers to a labeling method of setting (or attaching) a label (or label value) for the presence or absence of an error (or abnormality) at another specific time point (or another specific section) of the first robotics video. At this time, for the time point (or section) of the first robotics video where no label (or label value) is set by the additional selection labeling, a preset default label value (for example, an approval label) may be set.
[0520] That is, the server 200, in conjunction with the terminal 100, responds to the input of the user of the corresponding terminal 100 (or the user's selection / touch / control) for the first robotics video displayed on the corresponding terminal 100, and sets (or receives / inputs) the label (or label value) at another specific time point (or another specific section) of the corresponding first robotics video.
[0521] In addition, the server 200 receives one or more additional selection label values, one or more time-series segmentation selection label values, one or more body part-specific selection label values, a label value for sorting the order of the corresponding plurality of sub-robotics videos, identification information of the corresponding terminal 100, etc. related to the first robotics video transmitted from the terminal 100 at one or more other specific time points (or one or more other specific sections).
[0522] In an embodiment of the present invention, in the terminal 100, in response to a user's input, one or more additional selection label values are mainly described to be set (or received / input) at one or more other specific time points (or other specific intervals) of the corresponding first robotics video. However, the present invention is not limited thereto. The server 200 performs a video analysis function on the corresponding first robotics video and a comparison target video related to the corresponding first robotics video, and based on the execution result of the video analysis function, automatically sets one or more additional selection label values at one or more other specific time points (or other specific intervals) for the corresponding first robotics video.
[0523] Further, in the server 200, when one or more additional selection label values are set for the corresponding first robotics video at one or more other specific time points (or other specific intervals), the server 200 provides information regarding the one or more additional selection label values at one or more other specific time points (or other specific intervals) related to the set corresponding first robotics video to the terminal 100. In the corresponding terminal 100, the information regarding the one or more additional selection label values at one or more other specific time points (or other specific intervals) related to the corresponding first robotics video set by the server 200 is displayed, and in response to the input of the user of the corresponding terminal 100, it may be configured to determine whether to finally approve the one or more additional selection label values at one or more other specific time points (or other specific intervals).
[0524] At this time, before or after performing additional selection labeling on the corresponding first robotics video, the server 200, in conjunction with the terminal 100, performs additional hierarchical labeling on the corresponding one or more first robotics videos. Before or after performing the additional hierarchical labeling, additional selection labeling can also be performed on the corresponding first robotics video. Here, the additional hierarchical labeling (or additional hierarchical labeling) is user input feature engineering, which attaches a label (or label value) indicating the features of the corresponding first robotics video and divides (or classifies) the corresponding first robotics video into a plurality of sub-robotics videos according to the features.
[0525] That is, the server 200, in conjunction with the terminal 100, refers to (or based on) a plurality of preset label classifications related to the corresponding specific topic for the first robotics video displayed on the corresponding terminal 100, and according to the input of the user of the corresponding terminal 100 (or the user's selection / touch / control), sets (or receives / inputs) additional labels (or additional label values) at other specific time points (or other specific intervals) of the corresponding first robotics video.
[0526] In addition, the server 200 divides the first robotics video into a plurality of sub-robotics videos.
[0527] In the embodiments of the present invention, mainly described is that in the terminal 100, according to the user's input, one or more additional hierarchical labels (or additional hierarchical label values) are set (or received / input) at one or more other specific time points (or other specific intervals) of the corresponding first robotics video. However, it is not limited thereto. The server 200 performs a video analysis function on the corresponding first robotics video and the comparison target video related to the corresponding first robotics video, and based on the result of performing the video analysis function, one or more additional hierarchical label values can be automatically set respectively at one or more other specific time points (or other specific intervals) for the corresponding first robotics video.
[0528] Also, in the server 200, when one or more additional hierarchical label values are set for the corresponding first robotics video at one or more other specific time points (or other specific intervals), the server 200 provides the terminal 100 with information regarding the one or more additional hierarchical label values at one or more other specific time points (or other specific intervals) associated with the set corresponding first robotics video. At the corresponding terminal 100, information regarding the one or more additional hierarchical label values at one or more other specific time points (or other specific intervals) associated with the corresponding first robotics video set by the server 200 is displayed, and in response to the input of the user of the corresponding terminal 100, it may be configured to determine the presence or absence of final approval for the one or more additional hierarchical label values at the corresponding one or more other specific time points (or other specific intervals).
[0529] The second robotics video information 1912 is a video that is a predicted value of the second collective intelligence robotics 1906, and is defined as a predicted value of an advanced prediction model by repeated application of the repeated induction and / or inference algorithm 1200 in the present invention.
[0530] The second robotics video information 1912 has a second attribute 1226 and a second target attribute 1227 in FIGS. 4 to 6.
[0531] Also, the server 200 performs other machine learning on the artificial intelligence base based on information regarding the additionally selected and labeled first robotics video, etc., and generates (or confirms) a classification value for the corresponding first robotics video based on the results of the other machine learning. Here, the classification value for the corresponding first robotics video (or the classification value of the corresponding first robotics video) may be a value obtained by classifying additional selection labeling values, additional hierarchical labeling values, etc. by the same item.
[0532] That is, the server 200 performs other machine learning (or other artificial intelligence / other deep learning) using information about the additional selection-labeled first robotics video as an input value of the preset classification model, and generates (or confirms) a classification value for the corresponding first robotics video based on the result of the other machine learning (or the result of the other artificial intelligence / the result of the other deep learning).
[0533] The second robotics video information 1912 is labeled by the user using the second robotics selection labeling 1908 method, and the second robotics classification model 1909 is induced and / or inferred from the labeled data. The classified second robotics label information 1910 is input into the third collective intelligence robotics. The following is repeated.
[0534] The first robotics label information 1905 and the second basic robotics video information 1907 are learned as a single model of the second collective intelligence robotics 1906. The first collective intelligence robotics 1902 is programmed by the input of the first basic robotics video information 1901, and the second collective intelligence robotics 1906 is programmed by the input of the second basic robotics video information 1907 and the first robotics label information 1905. The following is repeated.
[0535] Further, the server 200 uses, as input values, the classification value for the generated corresponding first robotics video (or the classification value of the corresponding first robotics video), the information regarding the additionally selectively labeled first robotics video, the corresponding first robotics video, the meta information associated with the corresponding first robotics video, the comparison target video, the meta information associated with the corresponding comparison target video, etc., to perform other machine learning (or other artificial intelligence / other deep learning), and based on the result of the other machine learning (or the result of the other artificial intelligence / the result of the other deep learning), generates a second robotics video corresponding to the corresponding first robotics video. At this time, the second robotics video may be an operation-related video such as an avatar, an item, or a robot generated based on the first robotics video, a video in which the first robotics video is updated, or the like.
[0536] That is, the server 200 uses, as input values to the preset prediction model, the classification value for the generated corresponding first robotics video (or the classification value of the corresponding first robotics video), the information regarding the additionally selectively labeled first robotics video, the corresponding first robotics video, the meta information associated with the corresponding first robotics video, the comparison target video, the meta information associated with the corresponding comparison target video, etc., to perform other machine learning (or other artificial intelligence / other deep learning), and based on the result of the other machine learning (or the result of the other artificial intelligence / the result of the other deep learning), generates a second robotics video associated with the corresponding first robotics video.
[0537] Further, the server 200 transmits (or provides) the generated second robotics video to the terminal 100.
[0538] Referring to FIG. 19, from the perspective of the model, the second basic robotics video information 1907, which is the output and / or generated data of the metabus, and the first robotics label information 1905, which is the output and / or generated data of the first collective intelligence robotics 1902, may have different advantages and disadvantages in terms of accuracy and refinement. Although the two processes are labeling processes in different forms, in order for the model to accommodate only the advantages of the two processes, the modified second basic robotics video information 1907 of the approach for the two labels and the first robotics label information 1905, which is the output and / or generated data, are used as the training data of the same model (single model) rather than different models. After being initially labeled as the first robotics selection labeling 1903 and the first robotics classification model 1904, the first robotics label information 1905, which is the output and / or generated data, is secondarily labeled by the second robotics selection labeling 1908, and this process is continuously repeated. The data labeled in the past (the first robotics label information, 1905) comes to be repeatedly used together with other label data (the second basic robotics video information, 1907), and the data and / or similar label values once learned also continue to appear in each repeated training (epoch), and a process that goes through several experiments is required. Each epoch is divided into learning operation units (mini batch size) by the total number of accumulated unit labels (batch size) to conduct various experiments, and in the corresponding process, the label values of the collective intelligence are selected, averaged, and reflected in the model.
[0539] The robotics selection labeling 1810 of the collective intelligence robotics 1803 is in the same manner as the selection labeling 1604 such as avatar, human, robot, etc.
[0540] In one embodiment of the present invention, in the case of an automatic surgery and dental treatment robot with a limited range of movement and combination, it is not necessary to perform hierarchical clustering, and the correct and incorrect parts of the video information are carefully labeled one by one by the robotics selection labeling 1810. In the case of a humanoid robot or a vehicle robot with high degrees of freedom and / or inference (including, for example, a dancing robot, a soccer-playing robot, a bipedal walking robot, etc.), in FIG. 17 for hierarchical clustering, after dividing the robotics video information in digital units 3 1705 and / or digital units 4 1706 and / or digital units 5 1707 by time-series segmentation labeling 1701, selection labeling 1702 for each body part, etc., the robotics selection labeling 1810 is performed.
[0541] Further, the server 200 repeatedly performs the selection labeling process, the classification model inference process, the prediction model inference process, the additional selection labeling process for the generated first robotics video, the additional classification model inference process, and the additional prediction model inference process (for example, steps S2910 to S2980 above) for the operation-related videos of a plurality of actual humans, virtual avatars, items, etc. collected from a plurality of terminals 100 in relation to the corresponding specific topic, and generates (or updates) the second robotics video that has been collectively intellectualized in relation to the corresponding specific topic.
[0542] At this time, the server 200 may provide the second robotics video that was last updated (or newly generated) to the plurality of terminals 100 that provided the operation-related videos of actual humans, virtual avatars, items, etc. in real time or in response to a request from a specific terminal 100 in relation to the corresponding specific topic.
[0543] As a result, all terminals 100 or specific terminals 100 that have provided the server 200 with operation-related videos of actual humans, virtual avatars, items, etc. related to the corresponding specific topic can be provided with the latest collectively intelligent second robotics video in relation to the corresponding specific topic (or in relation to the comparison target video related to the corresponding specific topic).
[0544] The "first, second, third,... robotics video information 1813" repeatedly output and / or generated by the GAN and / or GNN robotics prediction model is repeatedly learned as a single model with the basic robotics video information 1802. Robotics selection labeling 1810 is repeatedly executed. The robotics classification model 1811 is repeatedly induced and / or inferred, and the GAN and / or GNN robotics prediction model is repeatedly induced and / or inferred. The GAN and / or GNN robotics prediction model is included in the collective intelligence robotics 1803.
[0545] In one embodiment of the present invention, hierarchical labeling, time-series segmentation selection labeling 1701, selection labeling by body part 1702, etc. in the same manner as the information processing of avatar movements are repeatedly executed.
[0546] In addition, the server 200 issues (or distributes) NFTs (non-fungible tokens) for the first video, second video, first robotics video, second robotics video, etc. generated based on the raw data, operation-related videos of avatars and / or items, etc. provided by the terminal 100 in conjunction with a blockchain server (not shown).
[0547] The NFT (or NFT content) issued by the server 200 is related to any digital art owned by the owner who has the ownership of providing the operation-related video of the load data, the avatar, and / or the item, and is content (or MR content / sensory content) generated corresponding to the relevant digital art (for example, including the first video, the second video, the first robotics video, the second robotics video, etc.). It may be in a state where an address indicating a digital file, a unique identification code (for example, including information related to asset information, creator, owner, etc.) and the like are inserted into the token for the original digital asset.
[0548] In addition, the server 200 is configured such that, in relation to the issued NFT, markers are displayed together on one side of the screen of the terminal 100 on which the first video, the second video, the first robotics video, the second robotics video, etc. are displayed.
[0549] Also, when a marker displayed on one side of the first video, the second video, the first robotics video, the second robotics video, etc. is selected by touching of the user of the terminal 100, the server 200 confirms the NFT corresponding to the selected marker, and information regarding the confirmed NFT (for example, including information related to asset information, creator, owner, etc.) is displayed on one side of the screen of the terminal 100 (or in a pop-up form on the screen on which the first video, the second video, the first robotics video, the second robotics video, etc. are displayed). At this time, the terminal 100 may display the information regarding the confirmed NFT in the form of virtual reality, augmented reality, mixed reality, extended reality, etc.
[0550] In addition, the server 200 provides a trading function (or a selling function / ownership transfer function) and the like for the NFT related to the issued first video, second video, first robotics video, second robotics video, etc.
[0551] That is, the video information (including, for example, the first video, the second video, the first robotics video, the second robotics video, etc.) is video information to which an NFT is attached, and the video information platform providing system to which the NFT is attached is a flywheel with a circular circulation structure in which users, participants, and companies generate profits, make money, and double the fun factor.
[0552] Referring to FIG. 20, the virtual avatar generation and / or output platform providing system using GAN and / or GNN is a flywheel as a platform in which users (users), participants (each influencer 2001 or individuals who promote their characters on SNS), and companies (advertisers and / or manufacturers) each generate profits, make money, and double the fun factor.
[0553] The video information to which an NFT is attached is the "first, second, third,... video information" repeatedly output and generated by the GAN and / or GNN prediction model 1605.
[0554] Referring to FIG. 20, the GNN and / or GAN prediction model 1605 operates on the server 200 in FIG. 1. The GNN and / or GAN prediction model 1605 utilizes the basic video information (the first basic video information 1501 and the second basic video information 1505 in FIG. 15) provided by the user and the influencer 2001 to generate or output NFT avatars and items to the marketing platform 2003. Companies and investors can own the profile NFTs and product NFTs of the influencer 2001 and utilize them for marketing and / or corporate publicity. The profile (including, for example, videos, photos, etc.) is the generated avatar, and the product is the item.
[0555] In one embodiment of the present invention, it is programmed to use the deepfake of the user and the influencer 2001 to advertise on the marketing platform 2003 and automatically register on the domestic and international NFT markets.
[0556] In one embodiment of the present invention, the marketing platform 2003 means any platform that can be marketed.
[0557] NFTs on the metabus are linked as a medium for digital twins of avatars and items with real-world owners, producers, advertisers, physical goods, etc.
[0558] In addition, by providing promotional expenses to participants and issuing NFTs for avatars and items to users, uniqueness is given and value is measured, and profits are generated by refunding the cost based on the value.
[0559] Referring to FIG. 14, the server 200 separately objectifies the human body and links information such as gender, age, body shape, and being of Asian descent with meta-information. Items (such as goods) are separately objectified and linked with meta-information. At this time, each avatar ID is linked with the user ID, item ID, and NFT ID.
[0560] Diverse real-world value and property information may be included in the form of metadata and NFT-ized, which can be bought, sold, and traded while ensuring uniqueness in the form of item NFTs. The platform ensures that the ownership of the corresponding NFT can be the right to use real-world value, and the usage breakdown and steps of the service are linked with the platform's database, and the NFT meta-information is updated and referenced.
[0561] In one embodiment of the present invention, the real-world value for NFT ownership includes, for example, the right to use a digital card that is a patient's avatar.
[0562] Referring to FIG. 20, the server 200 in FIG. 1 can generate products actually sold as items in the metabus and provide instructions so that real products can be purchased in reality.
[0563] In one embodiment of the present invention, an influencer 2001 who uses the service of the present invention can promote his / her avatar and the service of the present invention on an SNS on his / her network. The server 200 can obtain publicity-related content uploaded to an SNS channel on the network. The server 200 can analyze users flowing in via an SNS channel on the network and can calculate publicity costs to be provided to the SNS on the network based on the analyzed results. The server 200 can generate and provide different links for each influencer 2001, and can provide compensation to the influencer 2001 for users flowing in via the corresponding link. Further, the server 200 can also analyze whether a user is registered, the amount of item purchases, etc., and provide additional compensation to the influencer 2001.
[0564] In one embodiment of the present invention, the influencer 2001 includes entertainers, actors, athletes, etc.
[0565] In one embodiment of the present invention, NFTs are also assigned to each area including land, sea, and buildings in the metaverse, and they are made to play a role like a real estate register. Users trade each area using NFTs.
[0566] In one embodiment of the present invention, each object in the metaverse game may be composed of composite elements such as patterns, colors, materials, designs, etc. The server 200 NFTs in conjunction with meta information such as brand, product ID, seller ID, producer ID, advertiser ID, owner ID, etc. Further, the server 200 separately objectifies hats, accessories, and clothing, and each object is linked with meta information of a user, a producer, a uniqueness ID, or a representative object ID. At this time, each item ID may be linked with an NFT ID. Further, the server 200 assigns an NFT to an item purchased by a user such as an accessory, and configures it so that a transaction based on this is possible in the metaverse.
[0567] In one embodiment of the present invention, the server 200 provides dental, plastic surgery, and / or other store contents in the metabus, and if the cost of a desired procedure or operation and / or the cost of an item is paid, it uses GAN and / or GNN to change a certain part or the whole of the avatar and / or the digital cad. The server 200 issues NFTs to the digital cad synthesized with the purchased items (including, for example, surgical equipment, surgical instruments, surgical techniques, etc.) and the user's own avatar character. The user can be issued with NFTs for the corresponding digital cad and can sell them to obtain profits. That is, according to an embodiment of the present invention, while applying various combinations of items to the digital cad through GAN and / or GNN, it provides an element of fun, and by issuing NFTs to the synthesized digital cad, it provides uniqueness and can also obtain profits through this.
[0568] In one embodiment of the present invention, the server 200 issues NFTs to the avatar synthesized with the purchased items. The user can be issued with NFTs for the corresponding avatar and can sell them to obtain profits. That is, according to an embodiment of the present invention, while coordinating various combinations of items to the avatar through GAN and / or GNN1605, it provides an element of fun, and by issuing NFTs to the synthesized avatar, it provides uniqueness and can also obtain profits through this. Also, the server 200 configures to attach NFTs to the items purchased by the user such as accessories so that transactions based on this can be possible within the metabus.
[0569] In one embodiment of the present invention, the server 200 provides services such as trying on makeup, trying on clothes, receiving recommendations for makeup styles and fashion styles, substituting one's own face into the video of a celebrity, and checking the style.
[0570] In addition, the server 200 provides correction and execution stop alerts for minor mistakes and fatal mistakes made by the user based on label values (including, for example, an approval label corresponding to a correct one and a rejection label corresponding to an incorrect one) in response to user input in the labeling process (including, for example, a selection labeling process, a hierarchical labeling process, a time-series segmentation selection labeling process, a selection labeling process for each body part, etc.) for the load data, the first video, the second video, the operation-related video of the avatar and / or item, the first robotics video, the second robotics video, etc.
[0571] Referring to FIG. 16, the server 200 performs supervised learning on the artificial intelligence by selection labeling 1604 for correct and incorrect ones based on the user's judgment. Also, the server 200 intervenes with correction and execution stop alerts for minor mistakes and fatal mistakes made by the user on the terminal 100.
[0572] In an embodiment of the present invention, the video information is repeated as "second, third, fourth, …".
[0573] The video information is the repeated predicted values of the GAN and / or GNN prediction model 1605 and is "the first, second, third, … video information".
[0574] In one embodiment of the present invention, the collective intelligence robotics 1803 interacts with humans in a way that provides a warning signal (or alert signal). The automated surgical artificial intelligence that affects the patient's life is not a substitute for the doctor, but is included in the robotic arm steering device that assists the doctor in performing delicate surgery during the surgical process in a haptic concept. When attempting to perform an incorrect surgery, warning signals such as vibrations can be used to interact with and intervene with the doctor through zoom. If the corresponding warning signal is ignored and the surgery is performed, it may be used as separate label data stating that "it is correct to act in that way in the corresponding situation". Thus, while the artificial intelligence in the virtual world intervenes in and provides assistance for the doctor's surgery in the real world, the more users there are, the more refined it becomes due to its feedback.
[0575] In one embodiment of the present invention, for the actions of avatars, humans, robots, etc. in videos labeled with the not ACCEPT label or the REJECT label, the artificial intelligence learns with a teacher and sends an alert. The alert is also possible in virtual surgery, virtual driving, flying, etc. in a VR simulator, and is also possible in actual surgery, actual driving, flying, etc.
[0576] In one embodiment of the present invention, the alert is also possible through video information, audio information, haptic devices, etc.
[0577] In one embodiment of the present invention, if there is an error when a surgeon performs a gastric cancer surgery, the video is labeled with the REJECT label. The artificial intelligence learns with a teacher in response to this. When the artificial intelligence doctor robot assists in a gastric cancer surgery, it senses the doctor's incorrect surgical actions in the virtual surgery game and / or the actual gastric cancer surgery and sends an alert.
[0578] In one embodiment of the present invention, when the user adjusts a fighter plane in the fighter plane adjustment of a virtual war game and / or is shot down by an enemy plane, if the user attaches an ACCEPT or REJECT label to this video, the artificial intelligence performs supervised learning on this, senses an incorrect adjustment in the flight combat of an actual fighter plane pilot, and sends an alarm.
[0579] For example, in a virtual police game on a VR treadmill, when the user (in the role of a thief) steals an item or attaches a REJECT label to an act of committing a crime, the artificial intelligence performs supervised learning on this, senses the act of the thief in an actual security system, and sends an alarm.
[0580] Also, the step of transmitting information regarding the video information may be a step of performing a correction operation for a mistake made by the user or a step of the robot autonomously operating itself.
[0581] That is, the collective intelligence robotics 1803 in FIGS. 18 to 19 performs supervised learning on the visual data labeled by robotics selection labeling 1810, and the artificial intelligence robot operates itself. The terminal 100 performs a correction operation or an autonomous operation for a mistake made by the user.
[0582] In one embodiment of the present invention, the robotics video information is repeated as "second, third, fourth,...".
[0583] The video information is the repeated predicted values of a GAN and / or GNN robotics prediction model, and is "the first, second, third,... robotics video information".
[0584] In one embodiment of the present invention, the information with an ACCEPT label, a REJECT label, a not ACCEPT label, or a not REJECT label is used for the artificial intelligence to alarm the user, and is used for the artificial intelligence to operate itself to solve or avoid the problems that have occurred. The autonomous operation of the collective intelligence robotics 1803 is possible in a VR simulator and is also possible in actual reality.
[0585] In one embodiment of the present invention, it is also possible in virtual surgery, autonomous operation of various drones (VEHICLES), autonomous flight, or autonomous operation of humanoid robots.
[0586] In one embodiment of the present invention, an advanced surgical medical artificial intelligence corrects and intervenes with an Alert to abort the performance of minor or fatal mistakes made by a doctor while performing surgery (such as performing a procedure, treatment, etc.) on an artificial cadaver and an actual patient using a robotic arm, thereby providing assistance in real-world surgeries. The VR simulator is gamified in such a way that the artificial intelligence robotic arm operates and compensates for the labeling of surgical information by the operating doctor. Furthermore, for the labeled surgical information, the existing algorithm model is additionally fine-tuned to enhance the medical artificial intelligence. Ultimately, the artificial intelligence robotic arm proceeds with surgery on an actual human body, and in response, the doctor can perform the labeling.
[0587] In one embodiment of the present invention, an automatic surgical robot in a virtual surgery game can perform a gastric cancer surgery in a surgical VR simulator. If a doctor makes a selection and labeling 1604 in the virtual surgery on the simulator, the artificial intelligence performs supervised learning based on this, and the artificial intelligence doctor robot gradually becomes more advanced. The advanced artificial intelligence doctor robot can then automatically perform an actual surgery. The doctor makes a selection and labeling again, and the artificial intelligence becomes even more advanced. Through the repeated algorithm, the collective intelligence robotics 1803 becomes an autonomously operating artificial intelligence doctor robot or an artificial intelligence dentist robot.
[0588] In one embodiment of the present invention, in a virtual fighter plane flight game, when a user presses the ACCEPT button and attaches an ACCEPT label to a video in which a vehicle robot adjusts a fighter plane to shoot down an enemy plane, the artificial intelligence performs supervised learning based on this and learns the adjustment of the virtual or actual fighter plane by a fighter plane operator. It is possible to perform evasive maneuvers and attack maneuvers through proactive operation in a virtual fighter plane flight game or an actual fighter plane flight.
[0589] In one embodiment of the present invention, in virtual vehicle adjustment, if a real human makes a selection labeling 1604 for the autonomous driving of the VEHICLE robot on the VR simulator, the artificial intelligence performs supervised learning on this, and the VEHICLE robot becomes increasingly sophisticated. The sophisticated robot can then automatically perform actual driving, and if a real human makes a robotics selection labeling 1810 for this again, the artificial intelligence becomes even more sophisticated.
[0590] In one embodiment of the present invention, in a virtual dance competition on a VR treadmill, if an actual dance expert or domain expert or robotics developer or user makes a robotics selection labeling 1810 for a video in which a humanoid robot performs a dance competition, the artificial intelligence performs supervised learning on this, and the operation of the humanoid robot becomes increasingly sophisticated.
[0591] The information processing system 10 using the collective intelligence may further include an external server (not shown).
[0592] The external server can be connected to th...
Claims
1. A terminal that transmits one or more load data collected in relation to a specific topic, meta information related to the load data, a comparison target video, meta information related to the comparison target video, and identification information of a terminal, Receives one or more load data related to a specific topic transmitted from the terminal, meta information related to the load data, a comparison target video, meta information related to the comparison target video, and identification information of the terminal, Performs selection labeling on the one or more load data in conjunction with the terminal, Performs machine learning on an artificial intelligence basis based on information regarding the load data subjected to the selection labeling, generates a classification value for the load data based on the result of the machine learning, Performs machine learning using, as input values, the classification value for the generated load data, information regarding the load data subjected to the selection labeling, the load data, meta information related to the load data, the comparison target video, and meta information related to the comparison target video, and generates a first video corresponding to the load data based on the result of the machine learning, A server that transmits the generated first video to the terminal, and an information processing system using collective intelligence, including the server.
2. The server, Performs additional selection labeling on the first video in conjunction with the terminal, Performs machine learning on an artificial intelligence basis based on information regarding the first video subjected to the additional selection labeling, generates a classification value for the first video based on the result of the machine learning, Performs machine learning using, as input values, the classification value for the generated first video, information regarding the first video subjected to the additional selection labeling, the first video, meta information related to the first video, the comparison target video, and meta information related to the comparison target video, and generates a second video corresponding to the first video based on the result of the machine learning, The information processing system using collective intelligence according to claim 1, characterized in that the generated second video is transmitted to the terminal.
3. The server, In relation to the specific topic, for a plurality of raw data provided from a plurality of terminals, the selection labeling execution process, the classification value generation process, the first video generation process, the additional selection labeling execution process for the generated first video, the additional classification value generation process for the generated first video, and the second video generation process are each repeatedly performed, and a second video intellectualized in relation to the specific topic is generated. The information processing system using collective intelligence according to claim 1 is characterized by this.
4. The step of the server receiving one or more pieces of raw data related to a specific topic transmitted from a terminal, meta information related to the raw data, a comparison target video, meta information related to the comparison target video, and identification information of the terminal; The step of the server performing selection labeling on the one or more pieces of raw data in conjunction with the terminal; The step of the server performing machine learning on an artificial intelligence base based on information regarding the raw data subjected to the selection labeling, and generating a classification value for the raw data based on the result of the machine learning; The step of the server performing machine learning using, as input values, the classification value for the generated raw data, information regarding the raw data subjected to the selection labeling, the raw data, meta information related to the raw data, the comparison target video, and meta information related to the comparison target video, and generating a first video corresponding to the raw data based on the result of the machine learning; The step of the server transmitting the generated first video to the terminal; The step of the terminal outputting the first video transmitted from the server, an information processing method using collective intelligence including this.
5. The step of performing selection labeling on the one or more pieces of raw data is characterized by setting a label value at at least one of one or more specific time points and one or more specific intervals among the raw data according to user input for the raw data displayed on the terminal. The information processing method using collective intelligence according to claim 4.
6. The step of performing selection labeling on the one or more pieces of raw data For the load data displayed in the video display area of the terminal, according to the input of the user of the terminal, a label value is set for the correct or incorrect behavior of the movement of the object included in the load data at a specific time point or in a specific section. The information processing method using collective intelligence according to claim 4 is characterized by this.
7. Before or after the step of performing selective labeling on the one or more load data by the server, a step of performing hierarchical labeling on the one or more load data in conjunction with the terminal is further included. The information processing method using collective intelligence according to claim 4 is characterized by this.
8. The step of performing hierarchical labeling on the one or more load data is For the load data displayed on the terminal, based on a plurality of preset label classifications, according to the input of the user, a process of setting a label value at at least one of other specific time points and other specific sections of the load data; The information processing method using collective intelligence according to claim 7 is characterized by including a process of dividing the load data into a plurality of sub-load data.
9. The step of generating a classification value for the load data based on the result of the machine learning is Performing machine learning with the information related to the selectively labeled load data as the input value of a preset classification model, and generating a classification value for the load data based on the result of the machine learning. The information processing method using collective intelligence according to claim 4 is characterized by this.
10. The step of generating a first video corresponding to the load data based on the result of the machine learning is Performing machine learning with the generated classification value for the load data, the information related to the selectively labeled load data, the load data, the meta information related to the load data, the comparison target video, and the meta information related to the comparison target video as the input values of a preset prediction model, and generating a first video related to the load data based on the result of the machine learning. The information processing method using collective intelligence according to claim 4 is characterized by this.
11. The step of performing additional selective labeling on the first video in conjunction with the terminal by the server The server performs machine learning on the artificial intelligence base based on the information regarding the first video with additional selected labeling, and generates a classification value for the first video based on the result of the machine learning. The server performs machine learning using, as input values, the generated classification value for the first video, the information regarding the first video with additional selected labeling, the first video, the meta information associated with the first video, the comparison target video, and the meta information associated with the comparison target video, and generates a second video corresponding to the first video based on the result of the machine learning. The server transmits the generated second video to the terminal. The terminal outputs the second video transmitted from the server. The server repeatedly performs, for each of a plurality of raw data provided from a plurality of terminals in relation to the specific topic, the selection labeling process, the classification model inference process, the prediction model inference process, the additional selected labeling process for the generated first video, the additional classification model inference process, and the additional prediction model inference process, and generates a second video that is collectively intelligent in relation to the specific topic. The information processing method using collective intelligence according to claim 4 is further characterized by including this step.
12. The step of performing additional selected labeling on the first video The process in which the terminal divides the first video into a plurality of sub-videos based on the information regarding the sub-data divided into a plurality by performing the hierarchical labeling function on the raw data. The process in which, for the plurality of divided sub-videos, the terminal inputs a label value for a correct action or a label value for an incorrect action according to the user's input. The process in which the terminal inputs a label value indicating the order of the plurality of sub-videos according to the user's input in order to sort the order of the plurality of sub-videos. The process in which the terminal transmits to the server the label values for the correct and incorrect actions for the plurality of input sub-videos, the label value for sorting the order of the plurality of sub-videos, and the identification information of the terminal. The process of the server receiving, by performing the time-series segmentation selection labeling function for the first video, label values for correct and incorrect actions with respect to the plurality of sub-videos transmitted from the terminal, label values for arranging the order of the plurality of sub-videos, and identification information of the terminal, is included, and the information processing method using collective intelligence according to claim 11 is characterized by this.
13. The step of performing additional selection labeling for the first video includes: The process of the terminal dividing the first video into a plurality of sub-videos based on information regarding sub-load data divided into a plurality by performing the hierarchical labeling function for the load data; The process of the terminal receiving, respectively, label values for the operation order of the avatars included in the plurality of divided sub-videos; The process of the terminal receiving a label value indicating the order of the plurality of sub-videos in response to a user input in order to arrange the operation order by body part for the operations of the avatars included in the plurality of sub-videos; The process of the terminal transmitting the label values for the operation order of the avatars included in the plurality of input sub-videos, the label values for arranging the order of the plurality of sub-videos, and the identification information of the terminal to the server; The process of the server receiving, by performing the selection labeling function by body part for the first video, the label values for the operation order of the avatars included in the plurality of sub-videos transmitted from the terminal, the label values for arranging the order of the plurality of sub-videos, and the identification information of the terminal, and the information processing method using collective intelligence according to claim 11 is characterized by this.
14. Collect operation-related videos related to at least one of actual humans, avatars, and items in relation to a specific topic, and meta information related to the operation-related videos. To embody the collected operation-related videos as the operations of an actual robot, reconstruct the collected operation-related videos as robot operation videos, perform selection labeling on the robot operation videos in conjunction with a terminal, perform machine learning on an artificial intelligence foundation based on the information regarding the selected-labeled robot operation videos, generate a classification value for the robot operation videos based on the results of the machine learning, generate a first robotics video corresponding to the robot operation videos based on the generated classification value for the robot operation videos, the information regarding the selected-labeled robot operation videos, the robot operation videos, the meta information related to the robot operation videos, comparison target videos, and the meta information related to the comparison target videos, and a server that transmits the generated first robotics video to the terminal, A terminal that outputs the first robotics video transmitted from the server, an information processing system using collective intelligence including the same.
15. The server, In conjunction with the terminal, perform additional selection labeling on the first robotics video, perform machine learning on an artificial intelligence foundation based on the information regarding the additionally selected-labeled first robotics video, generate a classification value for the first robotics video based on the results of the machine learning, perform machine learning using, as input values, the generated classification value for the first robotics video, the information regarding the additionally selected-labeled first robotics video, the first robotics video, the meta information related to the first robotics video, comparison target videos, and the meta information related to the comparison target videos, generate a second robotics video corresponding to the first robotics video based on the results of the machine learning, and transmit the generated second robotics video to the terminal. The information processing system using collective intelligence according to claim 14, characterized in that.
16. The server, In relation to the specific topic, for the operation-related video associated with at least one of a plurality of actual humans, avatars, and items provided from a plurality of terminals, the selection labeling execution process, the classification value generation process, the first robotics video generation process, the additional selection labeling execution process for the generated first robotics video, the additional classification inference value generation process for the generated first robotics video, and the second robotics video generation process are each repeatedly performed, and a second robotics video that is group-intellectualized in relation to the specific topic is generated. The information processing system using group intelligence according to claim 15, characterized in that.
17. The step of collecting, by the server, an operation-related video associated with at least one of an actual human, an avatar, and an item in relation to a specific topic, and meta information related to the operation-related video; The step of reconstructing, by the server, the collected operation-related video as a robot operation video in order to embody the collected operation-related video as an operation of an actual robot; The step of performing selection labeling on the robot operation video in conjunction with the terminal by the server; The step of performing machine learning on an artificial intelligence base based on information regarding the robot operation video selected for labeling, and generating a classification value for the robot operation video based on the result of the machine learning by the server; The step of generating, by the server, a first robotics video corresponding to the robot operation video based on the classification value for the generated robot operation video, the information regarding the robot operation video selected for labeling, the robot operation video, the meta information related to the robot operation video, the comparison target video, and the meta information related to the comparison target video; The step of transmitting, by the server, the generated first robotics video to the terminal; An information processing method using group intelligence, including the step of outputting, by the terminal, the first robotics video transmitted from the server.
18. The information processing method using collective intelligence according to claim 17, further comprising, before or after the step of performing selective labeling on the robot operation video by the server, a step of performing hierarchical labeling on the robot operation video in conjunction with the terminal.
19. The step of performing additional selective labeling on the first robotics video by the server in conjunction with the terminal; The step of performing machine learning on the artificial intelligence base based on the information on the first robotics video subjected to the additional selective labeling by the server, and generating a classification value for the first robotics video based on the result of the machine learning; The step of performing machine learning by the server using, as input values, the generated classification value for the first robotics video, the information on the first robotics video subjected to the additional selective labeling, the first robotics video, the meta information related to the first robotics video, the comparison target video, and the meta information related to the comparison target video, and generating a second robotics video corresponding to the first robotics video based on the result of the machine learning; The step of transmitting the generated second robotics video to the terminal by the server; The step of outputting the second robotics video transmitted from the server by the terminal; The step of repeatedly performing, by the server, for the operation-related video related to at least one of a plurality of actual humans, avatars, and items provided from a plurality of terminals in relation to the specific topic, the selective labeling execution process, the classification value generation process, the first robotics video generation process, the additional selective labeling execution process for the generated first robotics video, the additional classification value generation process for the generated first robotics video, and the first robotics video generation process, and generating a second robotics video that is collectively intelligent in relation to the specific topic, the information processing method using collective intelligence according to claim 17.
Citation Information
Patent Citations
System and method for collective intelligence service
KR101859198B1