A server and its driving method for recommending user-customized models on a predictive model platform.
The server system addresses conventional model recommendation limitations by processing user data through sub-deep learning models to deliver context-aware, personalized recommendations using a predictive model, enhancing accuracy and reducing bias.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- NURIE AI INC
- Filing Date
- 2025-10-06
- Publication Date
- 2026-05-15
AI Technical Summary
Conventional model recommendation methods face challenges such as insufficient data for new users or items, inability to reflect current user interests, biased recommendations, lack of consideration for contextual factors like weather or emotional state, and limited support for voice, haptic, and biosignal data.
A server system that utilizes sub-deep learning models to process user data including text, image, and extraneous data to determine customized models, using a predictive model trained on multiple data types to provide accurate and context-aware recommendations.
Enables personalized model recommendations that reflect current user context and diverse data types, improving accuracy and reducing bias, while facilitating easy data analysis and model provision.
Smart Images

Figure 2026079729000001_ABST
Abstract
Description
[Technical Field]
[0001] Various embodiments of the present invention relate to a server that recommends user-customized models on a predictive model platform and a method for driving the same. [Background technology]
[0002] Recently, there has been a diverse range of interest in artificial intelligence technology, and as a result, learning about and utilizing AI is being done in a variety of ways by individuals.
[0003] However, conventional model recommendation methods have several drawbacks: insufficient data for new users or items makes appropriate recommendations difficult; recommendations based on past interactions cannot reflect changes in users' current interests; similar user-based recommendations can lead to bias within groups; recommendations based solely on user preferences can result in biased information exposure; they cannot reflect current context such as weather, location, or emotional state; how user data is collected and utilized is unclear; and existing models are mostly text / click-based, failing to properly reflect voice, haptic, image, and biosignals.
[0004] Therefore, there is a need for a system or method that can easily provide users with customized models while overcoming these problems. [Prior art documents] [Patent Documents]
[0005] [Patent Document 1] Korean Registered Patent Publication No. 10-2831702 (July 3, 2025) [Patent Document 2] Korean Registered Patent Publication No. 10-2823122 (June 17, 2025) [Overview of the project] [Problems that the invention aims to solve]
[0006] Therefore, this embodiment relates to a server and a method for driving the same that can easily provide a user with a customized model that has been input into a predictive model based on at least one data from among text data, image data, and / or extraneous data, based on the user's request. [Means for solving the problem]
[0007] According to various embodiments, a server that recommends a user-customized model on a predictive model platform includes a communication interface; a processor; the processor is configured to receive at least one user data from the user's external electronic device via the communication interface, input the at least one user data into a first sub-deep learning model, determine text data for the user, input the at least one user data into a second sub-deep learning model, determine image data for the user, input at least one of the text data and / or image data into the predictive model, determine a custom-made model for the user, and send the custom-made model to the user's external electronic device via the communication interface; the predictive model is learned based on a plurality of text data and a plurality of image data, a plurality of custom-made models, first result data in which the plurality of custom-made models are determined to be accurate, second result data in which some errors are determined to have occurred in the plurality of custom-made models, and third result data in which errors are determined to have occurred in the plurality of custom-made models.
[0008] According to various other embodiments, in a method for driving a server that recommends a user-customized model based on a prediction model, the method includes receiving at least one user data from an external electronic device of the user through a communication interface, inputting the at least one user data into a first sub-deep learning model through a processor to determine text data for the user, inputting the at least one user data into a second sub-deep learning model through the processor to determine image data for the user, inputting at least one or more of the text data and / or the image data into the prediction model through the processor to determine a customized model for the user, and being set to send the customized model to the external electronic device of the user through the communication interface. The prediction model is learned based on a plurality of text data and a plurality of image data, a plurality of customized models, first result data determined that the plurality of customized models are accurate, second result data determined that some errors have occurred in the plurality of customized models, and third result data determined that errors have occurred in the plurality of customized models.
Advantages of the Invention
[0009] The server according to this embodiment can preprocess at least one user data obtained from an external electronic device of the user and easily perform analysis on each data, input each analysis data into a prediction model, and extract a customized model by selection according to the output loss function, and has the advantage of being able to determine the suitability of the customized model and provide it to the user.
Brief Description of the Drawings
[0010] [Figure 1] A block diagram of an electronic device and a network according to various embodiments of the present invention is shown. [Figure 2]It is an exemplary diagram for explaining a method for a server to operate according to various embodiments of the present invention. [Figure 3a] It is an exemplary diagram for a server to perform a specific classification on text data according to various embodiments of the present invention. [Figure 3b] It is an exemplary diagram for a server to perform a specific classification on text data according to various embodiments of the present invention. [Figure 3c] It is an exemplary diagram for a server to perform a specific classification on text data according to various embodiments of the present invention.
Mode for Carrying Out the Invention
[0011] Hereinafter, various embodiments of this document will be described with reference to the accompanying drawings. The embodiments and the terms used therein are not intended to limit the technology described in this document to specific embodiments, but should be understood to include various modifications, equivalents, and / or alternatives of the corresponding embodiments. In connection with the description of the drawings, similar reference numerals can be used for similar components. Singular expressions can include plural expressions unless the context clearly indicates otherwise. In this document, expressions such as "A or B" or "at least one of A and / or B" can include all possible combinations of the listed items. Expressions such as "first," "second," "first," or "second" can modify the corresponding components in order or regardless of importance, and are only used to distinguish one component from another without limiting the corresponding component. When it is mentioned that a certain (e.g., first) component is "(functionally or communicatively) connected" or "connected" to a different (e.g., second) component, the certain component can be directly connected to the other component or can be connected through another component (e.g., a third component).
[0012] In this document, “configured to” can be used interchangeably with, depending on the context, “suitable for,” “capable of,” “modified to,” “made to,” “capable of,” or “designed to,” either in hardware or software terms. In some contexts, the expression “device configured to” may mean that the device “can” work with other devices or components. For example, the phrase “processor configured to perform A, B, and C” may mean a dedicated processor for performing the operations in question (e.g., an embedded processor), or a general-purpose processor (e.g., a CPU or application processor) that can perform the operations in question by running one or more software programs stored in a memory device.
[0013] The electronic devices described in the various embodiments of this document may include, for example, at least one of the following: a smartphone, a tablet PC, a desktop PC, a laptop PC, a netbook computer, a workstation, or a server.
[0014] Referring to Figure 1, various embodiments of a server 108 in a network environment 100 are described. The server 108 may include a bus 110, a processor 120, memory 130, an input / output interface 140, a display 150, and a communication interface 160. In some embodiments, the server 108 may omit at least one of its components or may have additional components. The bus 110 may include circuits that connect components 110-160 to each other and transmit communication (e.g., control messages or data) between the components. The processor 120 may include one or more of a central processing unit, an application processor, or a communication processor (CP). The processor 120 may, for example, perform calculations and data processing related to the control and / or communication of at least one other component of the server 108.
[0015] Memory 130 may include volatile and / or non-volatile memory. Memory 130 can store instructions or data related to at least one other component of server 108, for example. According to one embodiment, memory 130 can store software and / or programs.
[0016] The input / output interface 140 can, for example, transmit commands or data input from a patient or other external device to other components of the server 108, or output commands or data received from other components of the server 108 to a patient or other external device.
[0017] The display 150 may include, for example, a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic light-emitting diode (OLED) display, a microelectromechanical system (MEMS) display, or an electronic paper display. The display 150 can, for example, display various content to the patient (e.g., text, images, videos, icons, and / or symbols). The display 150 may include a touchscreen and can receive touch, gesture, proximity, or hovering input using, for example, an electronic pen or a part of the patient's body. The communication interface 160 can, for example, configure communication between the server 108 and an external device (e.g., a first external electronic device 102), a second external electronic device 104, or the server 108. For example, the communication interface 160 can be connected to the network 162 via wireless or wired communication to communicate with an external device (e.g., a second external electronic device 104 or the server 108).
[0018] Wireless communication may include cellular communication using at least one of the following: LTE, LTE-A (LTE Advance), CDMA (code division multiple access), WCDMA (wideband CDMA), UMTS (universal mobile telecommunications system), WiBro (Wireless Broadband), or GSM (Global System for Mobile Communications). In one embodiment, wireless communication may include at least one of the following: WiFi (wireless fidelity), Bluetooth (registered trademark), Bluetooth (registered trademark) Low Power (BLE), Zigbee (registered trademark), NFC (near field communication), Magnetic Secure Transmission, Radio Frequency (RF), or Body Area Network (BAN). In one embodiment, wireless communication may include GNSS. GNSS may include, for example, GPS (Global Positioning System), Glonass (Global Navigation Satellite System), Beidou Navigation Satellite System (hereinafter, "Beidou"), or Galileo, the European global satellite-based navigation system. Hereafter in this document, "GPS" may be used interchangeably with "GNSS". Wired communication may include at least one of the following: for example, USB (universal serial bus), HDMI (registered trademark) (high definition multimedia interface), RS-232 (recommended standard 232), power line communication, or POTS (plain old telephone service).Network 162 may include at least one of the following: a telecommunications network, such as a computer network (e.g., LAN or WAN), the Internet, or a telephone network.
[0019] According to various embodiments, the server 108 can operate an application consisting of multiple execution screens or a website consisting of multiple web pages, communicate with electronic devices (e.g., electronic devices 101, 102, 104, 106 in Figure 1) (e.g., smartphones, notebooks, etc.) via the network 161, process requests received from the electronic devices 101 regarding the application or web pages, and transmit the requested information to the electronic devices 101. The server 108 can transmit source code to the electronic devices 101 that enables each execution screen of the dedicated application or website to be displayed on the electronic devices 101, and the electronic devices 101 can receive the source code and display the execution screen requested by the user of the electronic devices 101 through the dedicated application or web browser. According to one embodiment, the configuration referred to as electronic device 101 in this disclosure may mean a user account connected to the platform provided by the server 108 through the electronic device.
[0020] Figure 2 is an illustrative diagram illustrating how the server operates according to various embodiments of the present invention.
[0021] Figures 3a, 3b, and 3c illustrate various embodiments of the present invention, illustrating how a server can perform specific classifications of text data.
[0022] In operation 201, the server 108 (e.g., processor 120 in Figure 1) can receive at least one user data from the user's external electronic devices 101, 102, 104, and 106 through a communication interface (e.g., communication interface 160 in Figure 1). According to one embodiment, the at least one user data may include user-generated text data, user-generated translation data, user-generated summary data, user-generated dialogue data, user-generated audio data perceived or emitted by the user, user-generated tactile data perceived or emitted by the user, user-generated olfactory data, user-confirmed or user-inputted input image data, and user analysis data and user mathematical calculation data. Specifically, the text data consists of data for general natural language text and can consist of simple data such as the user's self-introduction, the user's opinion, and the user's questions, as well as text data consisting of various sentences and / or paragraphs. Furthermore, the translation data consists of data resulting from direct translations from one language to another, and more specifically, it can consist of text data created by users translating English texts into Korean.
[0023] Furthermore, summary data can consist of text data containing many sentences, or data that summarizes many documents and / or texts in one to three columns in the initial data. Interactive data can consist of conversational record data exchanged through an app or application, or with at least one corresponding object from other users, chatbots, and / or voice assistants. Voice data can consist of the user's speech, voice, and voice responses, as well as the speech, voice, and voice response data of other users. Tactile data can consist of data acquired through the input / output interface of tactile feedback equipment (e.g., input / output interface 140 in Figure 1) as user responses or inputs to vibration, pressure, and temperature. Olfactory data can consist of responses to odors generated by the user or odors emitted externally, and / or sensing data from the olfactory sensing module. Input image data can consist of image data that the user visually perceives or uploads. Furthermore, the analysis data can consist of statistics, charts, and judgment data generated by the user through data analysis, and more specifically, it can consist of data such as accelerator calculation results, machine running model result interpretations, and personal reviews. In addition, the mathematical calculation data can consist of data directly related to mathematics, such as function values, calculator input values, and function graphs, as data for formulas, mathematical expressions, and calculation inputs performed by the user. Subsequently, server 108 can input at least one data from among text data, translation data, summary data, and dialogue data into the first sub-deep learning model, at least one data from among input image data, analysis data, and at least one user data into the second sub-deep learning model, and at least one data from among voice data, tactile data, and olfactory data into the third sub-deep learning model.
[0024] In operation 203, the server 108 (e.g., processor 120 in Figure 1) can input at least one user data into the first sub-deep learning model to determine text data for the user. According to one embodiment, the server 108 inputs at least one data from among document data, translation data, summary data, and dialogue data into the first sub-deep learning model to determine text data. If it is determined that the text data does not exceed a previously set text critical value, the text data can be classified into first text result data. If it is determined that the text data exceeds the text critical value, the text data can be classified into second text result data. Specifically, the server 108 can verify at least one data from among document data, translation data, summary data, and dialogue data. Thereafter, the server 108 can input at least one data from among the verified document data, translation data, summary data, and dialogue data into the first sub-deep learning model to determine text data for the user. Thereafter, server 108 can determine whether the text data is a first text result data consisting of data that performs simple tasks by comparing it with a previously set text critical value, or a second text result data consisting of data that requires complex tasks. The text critical value can be set as a criterion value for comparing whether the text data consists of simple data within one to three columns, such as "translate ~", "summarize ~", or "draw a picture like this". Thereafter, if the text data does not exceed the text critical value, server 108 can determine that the text data is a first text result data consisting of simple data within one to three columns, such as "translate ~", "summarize ~", or "draw a picture like this". If the text data exceeds the text critical value, server 108 can determine that the text data consists of a second text result data which is determined to contain at least four columns of data and / or a large amount of data even if it consists of one to three columns, such as multi-modal questions or complex requests.Thereafter, server 108 can input text data and at least one of the first text result data and / or second text result data into the prediction model.
[0025] In another embodiment, the first sub-deep learning model can be trained on multiple user data, multiple document data, multiple translation data, multiple summary data, and multiple dialogue data, along with multiple text data, multiple first text result data, and multiple second text result data. Specifically, the first sub-deep learning model for determining the text data for users can consist of a DNN-based FFN (feed-forward network), backpropagation, and RNN (recurrent neural network). An FFN is a neural network that progresses unidirectionally from the input layer to the hidden layer and from the hidden layer to the output layer. It is the simplest and most convenient, but it has the drawback of not being able to adjust the error weighting. In the case of backpropagation, the result value of the output layer is fed back to the input layer to reduce the error in the result, and the deviation of the weighting is often input, making it a DNN model that is relatively fast to learn and is often used. An RNN generates output nodes through intersection with hidden nodes via the context unit, but it does not affect the input and output layers and can be used in a timeline-like manner. Furthermore, the first sub-deep learning model applies the structure of a CNN, and examples of this are not limited.
[0026] Furthermore, although this embodiment demonstrated the ability to distinguish text data by inputting it into a first sub-deep learning model for analysis, it is possible to distinguish text data independently without the first sub-deep learning model.
[0027] In operation 205, the server 108 (e.g., processor 120 in Figure 1) can input at least one user data into the second sub-deep learning model to determine image data for the user. According to one embodiment, the server 108 inputs at least one data from the input image data, analysis data, and at least one user data into the second sub-deep learning model to determine image data. If it is determined that the image data does not exceed a previously set critical image value, the image data can be classified as first image result data. If it is determined that the image data exceeds the critical image value, the image data can be classified as second image result data. Specifically, the server 108 can verify the input image data and analysis data. Subsequently, the server 108 can input at least one data from the verified input image data and analysis data into the second sub-deep learning model to determine image data for the user. Subsequently, the server 108 can determine whether the image data is first image result data consisting of data that performs a simple operation by comparing it with a previously set critical image value, or second image result data consisting of data that requires a complex operation. Furthermore, the image critical value can be set by comparing whether the image data consists of one to three data points, or by setting a reference value of a certain time (e.g., 3 minutes) and / or a certain size (e.g., 10MB) at the time of input for video data. Thereafter, if the image data does not exceed the image critical value, the server 108 can determine whether the image data is simple image data consisting of one to three image data points, or whether the video data is within 3 minutes and / or 10MB or less in size, based on the first image result data. If the image data exceeds the image critical value, the server 108 can determine whether the image data consists of second result image data that compares whether the image data consists of at least four or more image data points and / or video data that is 3 minutes or longer and 10MB or more in size.Thereafter, server 108 can input at least one of the image data, first image result data, and / or second image result data into the prediction model.
[0028] In another embodiment, the second sub-deep learning model can be trained on multiple user data, multiple analysis data, and multiple input image data, as well as multiple image data, multiple first image result data, and multiple second image result data. Specifically, the second sub-deep learning model for determining the image data can utilize at least one of the following CNN structures: AlexNet, LENET-5, NIN, VGGNet, ResNet, WideResNet, GoogleNet, FractaNet, DenseNet, FitNet, RitResNet, HighwayNet, MobileNet, and DeeplySupervisedNet. More specifically, LeNet-5 is the most recent LeNet model, developed in the Yann LeCun lab in the 1990s, and can be used to recognize postal codes and numbers. The core point is that the LeNet structure does not differ significantly from current CNNs, and convolution and subsampling are used to connect feature maps in a fully-connected manner. Next, AlexNet was the winning model at ILSVRC 2012, and at that time, the AlexNet model is considered to have revolutionized deep learning. This is because AlexNet, with its CNN structure, can significantly reduce past top 5 errors. Subsequently, AlexNet began applying CNN techniques to image neural networks, and proceeds with this structure, but what is unusual here is that instead of applying many filters at once, it can be split into two parts and analysis can be carried out using two GPUs. ZFNet is almost identical in structure to AlexNet, and in fact, its performance is improved while only slightly changing the parameters used in AlexNet, and this allows us to check how the filters are learned in the intermediate layer in order to properly train a CNN.As another example, GoogLeNet, the winning model at LSVRC 2014, has a deeper depth and a thicker width compared to AlexNet, but the number of parameters has been significantly reduced. This is because GoogLeNet introduced the concept of an inception module, which was used in a module that emerged from the concept of whether it is possible to search for more information when performing a non-linear filtering operation, even though filtering is a linear operation. The structure of GoogLeNet has layers similar to a network within the network structure, and these can be formed in a Network In Network (NIN) structure and performed non-linearly. Furthermore, the second sub-deep learning model applies a CNN structure, and examples for this are not limited.
[0029] Furthermore, although this embodiment demonstrated the ability to distinguish image data by inputting it into a second sub-deep learning model for analysis, it is possible to distinguish image data independently without a second sub-deep learning model.
[0030] In operation 207, the server 108 (e.g., processor 120 in Figure 1) can input at least one user data into the third sub-deep learning model to determine the extra data for the user. According to one embodiment, the server 108 inputs at least one data from at least one user data, voice data, tactile data, and olfactory data into the third sub-deep learning model to determine the extra data for the user. If it is determined that the extra data does not exceed a previously set extra critical value, the extra data can be classified as first extra result data. If it is determined that the extra data exceeds the extra critical value, the extra data can be classified as second extra result data.
[0031] Specifically, server 108 can verify at least one data from among audio data, tactile data, and olfactory data. Subsequently, server 108 can input at least one data from the verified audio data, tactile data, and olfactory data into a third sub-deep learning model to determine extra data for the user. Subsequently, server 108 can compare the extra data with a previously set extra critical value to determine whether it is first extra result data consisting of data for performing simple tasks, or second extra result data consisting of data requiring complex tasks. The extra critical value is determined by comparing whether the audio data in the extra data is within one minute, or whether there is a certain odor (e.g., human body odor) at the time of input of the extra data, or by comparing a certain time (e.g., 5 seconds) and / or contact range (e.g., 1 cm) at the time of input of the tactile data. 2 The threshold value can be set to whether contact occurred within ). Thereafter, if the extra data does not exceed the extra critical value, the server 108 can determine if the extra data is simple extra data and audio data is data within 1 minute, or if the extra data is determined to be a human scent, or tactile data within 1 second and / or 1 cm. 2 If it is determined that contact occurred internally, it can be determined that it is the first extra result data. If the extra data exceeds the extra critical value, it can be determined that the extra data is complex extra data and the audio data is longer than 1 minute, or the extra data is not determined to be a human scent, or the tactile data is longer than 1 second and 1 cm. 2 If it is determined that contact has occurred beyond a certain width, it can be determined that the data consists of second extra result data. Thereafter, server 108 can input at least one of the extra data, first extra result data, and / or second extra result data into the prediction model.
[0032] In another embodiment, the third sub-deep learning model can be trained on multiple user data, multiple audio data, multiple tactile data, and multiple olfactory data, along with multiple extra data, multiple first extra result data, and multiple second extra result data. Specifically, the third sub-deep learning model for outputting at least one first audio feature data to determine the extra data can be broadly composed of a deep learning model that includes a Mel-spectrogram, MFCC (Mel Frequency Cepstral Coefficients), a Wav2Vec model, and a HuBERT model structure. Here, the third sub-deep learning model can include a Mel-spectrogram and MFCC, which are most useful for quantifying the characteristics of the audio data and converting them into an input form that the model can understand. First, the Mel-spectrogram used in the third sub-deep learning model can be based on a spectrogram obtained by converting the audio data in relation to time and frequency, and then converted to a Mel scale that matches the user's auditory perception. The Mel-spectrogram processing method divides continuous data into small segments (frames) at regular time intervals, allowing for the division of at least one audio data point into each segment. Each divided frame can be processed in a manner similar to a Hamming window and flexibly concatenated. A Short-Time Fourier Transform (STFT) is then performed on each frame to convert the time-domain data into the frequency domain, thereby outputting it in spectrogram form. Subsequently, since the Mel-spectrogram is sensitive to low frequencies but insufficiently sensitive to high frequencies, a Mel-scale filter can be used to non-linearly transform the frequency axis of the spectrogram. The amplitude values of the spectrogram can be converted to a logarithmic scale, and the frequency band can be readjusted based on a logarithmic scale that reflects the human auditory perception of magnitude. Finally, a Mel-spectrogram in 2D data form containing time and Mel frequency components can be generated.The characteristics of such a Mel-spectrogram include the ability to provide high-resolution frequency information and the advantage of being usable for visual representation of at least one audio data set. Furthermore, MFCC, used in the third sub-deep learning model, is a variation of the Mel-spectrogram that can further compress the frequency information of audio data and extract small feature vectors suitable for model input. The MFCC processing process involves summarizing the logarithmic values of the Mel-filtered spectrogram based on energy density to highlight the most important information (low-frequency components) and remove unnecessary high frequencies, converting it into a Discrete Cosine Transform (DCT). From the DCT result, several of the most important coefficients (generally 12-13) can be selected to reduce the data size while preserving the basic sound quality of the audio data. Subsequently, the MFCC can capture even more dynamic information by adding changes over time (delta) and changes in the rate of change (delta-delta). The characteristics of such an MFCC include small data size, low computational cost, advantages for real-time processing, and the ability to well preserve the core acoustic features of audio data. Furthermore, the third sub-deep learning model can include the Wav2Vec model and the HuBERT model. More specifically, the Wav2Vec model and the HuBERT model included in the third sub-deep learning model are deep learning foundational models that play an important role in natural language processing and speech data recognition technologies for processing and analyzing speech data. Both models can learn useful representations from speech data and be utilized in a variety of speech data processing tasks.Furthermore, the Wav2Vec model included in the third sub-deep learning model configuration was developed at Facebook AI Research (FAIR) and can receive audio data as input and convert it into continuous latent representations. It can extract features from the audio data using a Convolutional Neural Network (CNN) and learn linguistic patterns from the extracted features using a Transformer. In addition, the Wav2Vec model included in the third sub-deep learning model configuration can use a self-supervised learning method, allowing it to learn to predict parts of the audio data that are hidden, and has the advantage of enabling efficient learning by utilizing large amounts of unlabeled audio data.
[0033] Furthermore, the Wav2Vec model included in the third sub-deep learning model configuration can demonstrate excellent performance even with small amounts of labeled data, enabling multilingual speech data recognition by learning speech data from various languages. It can analyze the emotions of speech data or extract specific timbre and tone features, and can be used for transfer learning in various tasks such as speech data generation and speaker identification. Additionally, the HuBERT model included in the third sub-deep learning model configuration was developed by Facebook AI Research as a successor to Wav2Vec. Similar to Wav2Vec, it is used with CNN and Transformer structures, and after extracting features from speech data, it can learn linguistic information using a Transformer. Specifically, as its name Hidden Unit BERT (HuBERT) suggests, the HuBERT model included in the third sub-deep learning model configuration can use BERT-style self-supervised learning, clustering speech data to generate latent labels (hidden units), and has the advantage of enabling more sophisticated speech data representation learning by training the model based on the clustered labels. Furthermore, the HuBERT model included in the third sub-deep learning model configuration differs from previous models in that it uniquely generates arbitrary labels by clustering the signals themselves during the learning process, allowing it to progressively learn better representations through many stages. In addition, the HuBERT model included in the third sub-deep learning model configuration can learn richer and more sophisticated speech data representations than Wav2Vec, can receive speech data input and translate it into other languages or convert it to text, can perform multimodal tasks by integrating speech data, olfactory data, and / or tactile data, and can maximize performance by utilizing a small amount of label data in data-scarce environments. In conclusion, the server 108 in this embodiment can easily output excess data by inputting at least one speech data into the third sub-deep learning model configured with various speech data conversion models.From this point onward, server 108 can input extraneous data into the predictive model.
[0034] Furthermore, although this embodiment demonstrated that extraneous data was analyzed by inputting it into a third sub-deep learning model, it is possible to distinguish extraneous data (inherently) without a third sub-deep learning model.
[0035] In operation 209, the server 108 (e.g., processor 120 in Figure 1) can input at least one of the following data—text data, image data, and / or extraneous data—into a predictive model to determine a customized model for the user. According to one embodiment, the predictive model can be trained on multiple text data and multiple image data, multiple customized models, first result data in which multiple customized models are judged to be accurate, second result data in which multiple customized models are judged to have some errors, and third result data in which multiple customized models are judged to have errors. The predictive model can also be trained on multiple extraneous data, multiple first text result data, multiple second text result data, multiple first image result data, multiple second image result data, multiple first extraneous result data, and multiple second extraneous result data, as well as multiple overall loss functions, multiple quality loss functions, multiple cost loss functions, and multiple efficiency loss functions. Specifically, the predictive model for determining a customized model for the user can be driven by an AI neural network model trained through unsupervised learning on the underlying data. The predictive model can be configured with increased ease of data collection and diverse data output values. As a structure for outputting image data from text data, the predictive model can utilize at least one of the following: BigScience's bloom and T0pp, EleutherAI's GPT series, Tsinghua UNIV's GLM series, Google's UL and T5 series, and META AI's OPT series. In one embodiment, the predictive model can be used in multiple crowdfunding models and implemented using the structures of Microsoft's Chat GPT, Google's BARD series, and NVIDIA's translation service platform transformer model. For example, the predictive model can be trained on a multimodal platform model based on at least one of the following data: text data, image data, and audio data, to implement diverse image generation.In one embodiment, this method can reduce the amount of labeled training data per task compared to existing deep learning methods. Once completed, diverse training is possible with a small amount of training data, making data collection and labeling easier and improving accuracy. Furthermore, the server 108 can perform the learning process of a prediction model for outputting state prediction data by using a prediction model to which arbitrary weight values are assigned to obtain result values (output data), comparing the obtained result values with the labeling data of the learning data, performing backpropagation based on the error, and optimizing the weight values. Specifically, learning the prediction model means training the prediction model based on the learning data and labeling data or unlabeled data, so that the prediction model determines the output data for the input data. In other words, the prediction model forms rules and makes judgments on the data. According to one embodiment, the server 108 can use multiple learning algorithms from among multiple learning algorithms for calculating the predicted values. For example, the ensemble method can be used in the prediction model, and even better prediction performance can be obtained compared to using separate learning algorithms. Training a predictive model can be interpreted as adjusting the weights that the model possesses. In one example, various learning methods can be used, including supervised learning, unsupervised learning, reinforcement learning, imitation learning, and federated learning.
[0036] Although not shown in the diagram, server 108 may include an evaluation stage in the training process of the prediction model to evaluate the performance of the prediction model. In the evaluation stage, the prediction model may be evaluated using an evaluation dataset. The evaluation of the prediction model may be a stage in which the prediction model trained in the training stage is evaluated and predictions are made on new data using the prediction model. Specifically, the evaluation stage may be a stage in which the trained prediction model can be generalized to new data.
[0037] Furthermore, the predictive model is a type of artificial neural network (MS) with a long short-term memory (LSTM) structure, specifically designed to learn patterns from time-series and sequential data. It is primarily an improved version of the recurrent neural network (RNN), which predicts the future based on past information but suffers from the gradient vanishing problem, making it difficult to learn long sequences and handle long-term dependencies. The predictive model (LSTM) addresses this problem by incorporating a special gate structure to control the flow of information, enabling it to learn long-term dependencies. The predictive model operates by selectively remembering and forgetting information using three gates: an input gate, a forget gate, and an output gate. These gates can perform the following actions at each point in time: First, the input gate determines how much new information to accept; the forget gate determines how quickly previously remembered information is forgotten; and the output gate determines the amount of information to transmit next based on the current state. Based on the three gates at the top, the predictive model learns and adjusts itself to determine whether past information is necessary for the current prediction, and learns through learning whether information at the beginning of a sentence is necessary at the end of the sentence. Furthermore, the predictive model (LSTM) can be used in natural language processing such as text generation, translation, and sentiment analysis. In particular, it has the strength to grasp the context of the entire sentence in machine translation, it can process audio data in time series to recognize specific words and syntax, it can be used to predict future stock prices by learning past stock price fluctuation patterns, it can be used to predict behavior and scene classification by learning changes between frames over time in videos, and it can be used to make long-term predictions, such as predicting the probability of developing a disease, by analyzing a user's past medical history and medical records. In addition, the predictive model (LSTM) has the advantage of being able to remember and utilize even long periods of past data.However, the prediction model (LSTM) has high computational costs, and as the amount of data increases, the learning time can become long. It can include structures that are overcome and selectively used in many situations as other sequential models such as GRU (Gated Recurrent Unit) and Transformer models emerge.
[0038] According to another embodiment, as shown in FIGS. 3a, 3b, and 3c, the server 108 can input text data into a prediction model based on the following [Equation 1] to generate a first quality loss function, a first cost loss function, and a first efficiency loss function for the text data, and a first overall loss function, respectively.
[0039] [Equation 1] JPEG2026079729000002.jpg46144
[0040] Here, L is the first overall loss function, L quality is the first quality loss function, L cost is the first cost loss function, L efficiency is the first efficiency loss function, w1, w2, w3 are the respective weighting values for the first quality loss function, the first cost loss function, and the first efficiency loss function, L accuracy is the value obtained by dividing the correct score for the domain, L completeness is the completion score value for the domain, β1, β2 are the weighting values for calculating the first quality loss function, L token is the calculated value for the price comparison token, L overhead is the total value of other processing costs excluding the price, δ1, δ2 are the weighting values for calculating the first cost loss function, L time is the ratio value for the actual processing time, L cmputation is the ratio value for the number of tokens used, y1, y2 are the weighting values for calculating the first efficiency loss function.
[0041] Specifically, if server 108 confirms whether only text data is input to the prediction model from text data, image data, and extraneous data, or whether only text data is input to the prediction model from multiple data sets, it can output a first quality loss function, which is a quality loss function for text data, a first cost loss function, which is a cost loss function for text data, and a first efficiency loss function, which is an efficiency loss function for text data, so that the output customized model can be accurately provided. Furthermore, the first quality loss function can be input into the prediction model to check whether it is similar to or better than the first text result data or the second text result data, the first cost loss function can be compared in the prediction model to the first text result data or the second text result data in terms of cost increase or decrease, and the first efficiency loss function can be input into the prediction model to check whether it is the most efficient in comparison with the second text result data in terms of final response value output time. Furthermore, Server 108 can be identified as the best user-customized model if it shows 95% similarity in the first quality loss function output by the prediction model, is 90% or more cheaper compared to other models in the first cost loss function, and is equal in the comparison of total required time in the first efficiency loss function. Additionally, Server 108 can output a model that is a customizable model with some errors but is still acceptable if it shows 90% similarity in the first quality loss function output by the prediction model, is 50% or more cheaper compared to other models in the first cost loss function, and shows a 10% error in the comparison of total required time in the first efficiency loss function.
[0042] In another embodiment, as shown in Figures 3a, 3b, and 3c, the server 108 can input image data into a prediction model based on the following [Mathematical Equation 2] to generate a second quality loss function, a second cost loss function, a second efficiency loss function, and a second overall loss function for the image data.
[0043] [Mathematical formula 2] JPEG2026079729000003.jpg45144 Here, L is the second overall loss function, L quality This is the second quality loss function, L cost This is the second cost loss function, L efficiency is the second efficiency loss function, where w1, w2, and w3 are the weighted values for the second quality loss function, the second cost loss function, and the second efficiency loss function, respectively, L accuracy L is a value obtained by dividing the correct score for the domain. completeness is the completion score for the domain, β1 and β2 are weighted values for calculating the second quality loss function, and L token This is a calculated value relative to the price of the token, L overhead is the sum of the second cost of other processing excluding price, and δ1 and δ2 are weighted values for calculating the second cost loss function, L time This is a ratio value to the actual processing time, L cmputation is a ratio value to the number of tokens used, and y1 and y2 are weights used to calculate the second efficiency loss function.
[0044] Specifically, if server 108 confirms whether only image data is input to the prediction model from text data, image data, and extraneous data, or whether image data is input to the prediction model from multiple data sets, it can output a second quality loss function, which is a quality loss function for image data, a second cost loss function, which is a cost loss function for image data, and a second efficiency loss function, which is an efficiency loss function for image data, so that the output customized model can be accurately provided. Furthermore, the second quality loss function can be input into the prediction model to check whether it is similar to or better than the first image result data or the second image result data, the second cost loss function can be compared in the prediction model to the first image result data or the second image result data in terms of cost increase or decrease, and the second efficiency loss function can consist of values that can be input into the prediction model to check whether the second image result data is the most efficient in terms of comparison of final response value output time. Furthermore, Server 108 can be identified as the best user-customized model if it shows 95% similarity in the second quality loss function output by the prediction model, is 90% or more cheaper compared to other models in the second cost loss function, and is equal in the comparison of total required time in the second efficiency loss function. Additionally, if Server 108 shows 90% similarity in the second quality loss function output by the prediction model, is 50% or more cheaper compared to other models in the second cost loss function, and shows a 10% error in the comparison of total required time in the second efficiency loss function, it can be output as a customizable model that has some errors but is still acceptable.
[0045] In another embodiment, as shown in Figures 3a, 3b, and 3c, the server 108 can input the extraneous data into a prediction model based on the following [Mathematical Formula 3] to generate a third quality loss function, a third cost loss function, a third efficiency loss function, and a third overall loss function for the extraneous data, respectively.
[0046] [Mathematical formula 3] JPEG2026079729000004.jpg43144 Here, L is the third overall loss function, L quality This is the third quality loss function, L cost This is the third cost loss function, L efficiency is the third efficiency loss function, and w1, w2, w3 are the weighted values for the third quality loss function, the third cost loss function, and the third efficiency loss function, respectively, L accuracy L is a value obtained by dividing the correct score for the domain. completeness is the completion score for the domain, β1 and β2 are weights for calculating the third quality loss function, and L token This is a calculated value relative to the price of the token, L overhead is the sum of the third cost of other processing excluding price, and δ1 and δ2 are weighted values for calculating the third cost loss function, L time This is a ratio value to the actual processing time, L cmputation is a ratio value to the number of tokens used, and y1 and y2 are weights used to calculate the third efficiency loss function.
[0047] Specifically, if server 108 confirms that only the extraneous data is input into the prediction model from text data, image data, and extraneous data, or that the extraneous data is input into the prediction model from multiple data sets, it can output a third quality loss function, which is a quality loss function for the extraneous data, a third cost loss function, which is a cost loss function for the extraneous data, and a third efficiency loss function, which is an efficiency loss function for the extraneous data, so that the output customized model accurately provides this. Furthermore, the third quality loss function can be input into the prediction model to verify whether it is similar to or better than the first extraneous result data or the second extraneous result data, the third cost loss function can be compared in the prediction model to the first extraneous result data or the second extraneous result data in terms of cost increase or decrease, and the third efficiency loss function can consist of values that can be input into the prediction model to verify whether it is the most efficient in comparison with the final response value output time of the second extraneous result data. Furthermore, if Server 108 shows 95% similarity in the third quality loss function output by the prediction model, is 90% or more cheaper compared to other models in the third cost loss function, and is equal in the overall required time comparison in the third efficiency loss function, it can be determined to be the best user-customized model. Additionally, if Server 108 shows 90% similarity in the third quality loss function output by the prediction model, is 50% or more cheaper compared to other models in the third cost loss function, and shows a 10% error in the overall required time comparison in the third efficiency loss function, it can be output as a custom-made model that has some errors but is still acceptable.
[0048] In operation 211, the server 108 (e.g., processor 120 in Figure 1) can determine whether the customized model exceeds a pre-set first critical value. In one embodiment, the server 108 can output quality data, cost data, and efficiency data, each between 0 and 1, to confirm whether the user customized model determined by the prediction model exceeds a pre-set first critical value. In this embodiment, the server 108 can be set to a threshold value that prioritizes whether the customized model is functionally efficient or cost-sensitive when performing the routing technique. Furthermore, the first critical value can be set according to the administrator's settings: if functional efficiency is prioritized, the quality is 98% or more similar to a single model, costs are reduced by 70% or more, and time is reduced by 20%; if cost is prioritized, the first critical value can be set according to the user's convenience, with quality being 90% or more similar to a single model, costs being reduced by 85% or more, and time being reduced by 10% to 0%. Subsequently, server 108 can check whether it is functional or cost-effective and compare quality data, cost data, and efficiency data, respectively, using the first critical value set according to the relevant function.
[0049] In operation 213, the server 108 (e.g., processor 120 in Figure 1) can determine that a custom-made model is a user-customized model if it determines that the custom-made model does not exceed a first critical value. According to one embodiment, if the server 108 prioritizes functionality, it can determine that the functionality-prioritized first critical value (e.g., quality data = 98%, cost data = 70%, efficiency data = -20%) is not exceeded if it confirms that the same percentage value of quality data is 99.1%, the cost savings value of cost data is 78.4%, and the time required of efficiency data has increased to -24.1%. In another embodiment, if cost is the primary consideration, server 108 can determine that the first critical value prioritizing cost (e.g., quality data = 90%, cost data = 85%, efficiency data = 0%) is not exceeded if it confirms that the same percentage value for quality data is 91.6%, the cost reduction value for cost data is 87.9%, and the time required for efficiency data has increased by -15.1%. Subsequently, server 108 can determine that the customized model determined by the predictive model is the most suitable model for the user.
[0050] On the other hand, in operation 211, if the server 108 (e.g., processor 120 in Figure 1) determines that the customized model has exceeded the first critical value, then in operation 215, the server 108 (e.g., processor 120 in Figure 1) can compare the customized model to a second critical value set to be larger than the first critical value. According to one embodiment, if the server 108 prioritizes functionality, and confirms that the same percentage value of quality data is 95.7%, the cost savings value of cost data is 76.4%, and the time required for efficiency data has increased to -21.1%, then it can determine that the cost data and efficiency data do not exceed the functionality-prioritized first critical value (e.g., quality data = 98%, cost data = 70%, efficiency data = -20%), but the quality data does. In another embodiment, if the server 108 prioritizes cost, and confirms that the same percentage value for quality data is 90.6%, the cost reduction value for cost data is 79.2%, and the time required for efficiency data has increased to -15.1%, it can determine that the quality data and efficiency data do not exceed the first critical value prioritizing cost (e.g., quality data = 90%, cost data = 85%, efficiency data = 0%), but the cost data does. Subsequently, the server 108 can compare the quality data, cost data, and efficiency data for the customized model to a second critical value set higher than the first critical value. Furthermore, the second critical value can be set according to the administrator's settings: if functional efficiency is prioritized, the quality is 95% or more similar to a single model, costs are reduced by 60% or more, and time is reduced by 10%; if cost is prioritized, the second critical value can be set to 80% or more similar to a single model, costs are reduced by 70% or more, and time is reduced by 0-10% according to the user's convenience. Subsequently, server 108 can check whether it is functional or cost-effective and compare quality data, cost data, and efficiency data, respectively, using the second critical value set according to the relevant function.
[0051] In operation 217, if the server 108 (e.g., processor 120 in Figure 1) determines that the customized model has exceeded the first critical value but has not exceeded the second critical value, it can determine that some errors have occurred in the customized model. According to one embodiment, if the server 108 prioritizes functionality, and confirms that the same percentage value of the quality data is 95.7%, the cost savings value of the cost data is 76.4%, and the time required for the efficiency data has increased to -11.1%, it can determine that the cost data does not exceed the first critical value prioritizing functionality (e.g., quality data = 98%, cost data = 70%, efficiency data = -20%), but that the quality data and efficiency data have exceeded the first critical value, and that the second critical value prioritizing functionality (e.g., quality data = 95%, cost data = 60%, efficiency data = -10%) is not exceeded. In another embodiment, if server 108 determines that cost is the priority, and confirms that the same percentage value for quality data is 89.6%, the cost savings value for cost data is 79.2%, and the time required for efficiency data has increased to -5.1%, it can determine that the efficiency data does not exceed the first critical value for cost priority (e.g., quality data = 90%, cost data = 85%, efficiency data = 0%), but it can also determine that the quality data and cost data have exceeded the first critical value, and therefore the second critical value for cost priority (e.g., quality data = 80%, cost data = 70%, efficiency data = 10%) is not exceeded. Subsequently, server 108 can determine that the customized model determined by the prediction model contains some errors for the user or has slightly reduced functionality, but is still a usable model for the user.
[0052] On the other hand, in operation 215, if the server 108 (e.g., processor 120 in Figure 1) determines that the customized model has exceeded the first critical value and the second critical value, then in operation 219, the server 108 (e.g., processor 120 in Figure 1) can determine that the error in the customized model is very serious. According to one embodiment, if the server 108 prioritizes functionality, and confirms that the same percentage value of quality data is 95.7%, the cost savings value of cost data is 56.4%, and the time required of efficiency data has increased to 0.24%, then it can determine that the overall value has been exceeded at the first critical value prioritizing functionality (e.g., quality data = 98%, cost data = 70%, efficiency data = -20%), and that the quality data has not been exceeded at the second critical value prioritizing functionality (e.g., quality data = 95%, cost data = 60%, efficiency data = -10%), but the cost data and efficiency data have exceeded the second critical value. In another embodiment, if server 108 determines that cost is the priority, and confirms that the same percentage value of quality data is 74.6%, the cost-saving value of cost data is 71.2%, and the time required value of efficiency data has increased to 0.89%, it can determine that the overall data has exceeded the first critical value for cost priority (e.g., quality data = 90%, cost data = 85%, efficiency data = 0%), and that the cost data and efficiency data have not exceeded the second critical value for cost priority (e.g., quality data = 80%, cost data = 70%, efficiency data = 10%), but the quality data has exceeded the second critical value. Thereafter, server 108 can determine that the errors in the customized models that have exceeded the first critical value and the second critical value are very serious, and can return to operation 203 only for customized models that are above the first critical value and the second critical value, and repeatedly run the process until the predictive model does not exceed the first critical value and / or the second critical value.
[0053] In operation 221, the server 108 (e.g., processor 120 in Figure 1) can send only the customized models determined in operations 213 and / or 217 to the user's external electronic devices 101, 102, 104, and 106 via the communication interface (e.g., communication interface 160 in Figure 1). According to one embodiment, the server 108 can provide or recommend customized models that do not exceed the first critical value and / or the second critical value to the user's external electronic devices 101, 102, 104, and 106. Thereafter, the user can confirm on the display of their external electronic devices 101, 102, 104, and 106 that they are using or plan to use the customized models.
[0054] The server 108 in this embodiment has the advantage of being able to easily perform analysis on each data by pre-processing at least one user data obtained from the user's external electronic devices 101, 102, 104, and 106, inputting each analysis data into a predictive model, extracting a customized model based on the output loss function, determining the suitability of the customized model, and providing it to the user.
[0055] According to various embodiments, a server that recommends a user-customized model on a predictive model platform includes a communication interface; a processor; the processor is configured to receive at least one user data from the user's external electronic device via the communication interface, input the at least one user data into a first sub-deep learning model to determine text data for the user, input the at least one user data into a second sub-deep learning model to determine image data for the user, input at least one of the text data and / or image data into the predictive model to determine a user-customized model; and send the user-customized model to the user's external electronic device via the communication interface; the predictive model is learned based on a plurality of text data and a plurality of image data, a plurality of user-customized models, first result data in which the plurality of user-customized models are determined to be accurate, second result data in which some errors are determined to have occurred in the plurality of user-customized models, and third result data in which errors are determined to have occurred in the plurality of user-customized models.
[0056] According to various embodiments, the at least one user data includes text data created by the user, translation data created by the user, summary data created by the user, dialogue data created by the user, audio data perceived or emitted by the user, tactile data perceived or emitted by the user, olfactory data perceived or emitted by the user, input image data confirmed by the user's eyes or entered by the user, and analysis data of the user and mathematical calculation data of the user.
[0057] According to various embodiments, the processor is configured to input at least one data from the text data, translation data, summary data, and dialogue data into the first sub-deep learning model, determine the text data, classify the text data with first text result data if it is determined that the text data does not exceed a previously set text critical value, and classify the text data with second text result data if it is determined that the text data exceeds the text critical value, and the first sub-deep learning model is trained on multiple user data, multiple text data, multiple translation data, multiple summary data, and multiple dialogue data, and multiple text data, multiple first text result data, and multiple second text result data.
[0058] According to various embodiments, the processor is configured to input at least one data from the input image data, the analysis data, and the at least one user data to the second sub-deep learning model to determine the image data, classify the image data by first image result data if it is determined that the image data does not exceed a previously set critical image value, and classify the image data by second image result data if it is determined that the image data exceeds the critical image value, and the second sub-deep learning model is trained on multiple user data, multiple analysis data, and multiple input image data, and multiple image data, multiple first image result data, and multiple second image result data.
[0059] According to various embodiments, the processor is configured to input at least one data from the at least one user data, the voice data, the tactile data, and the olfactory data to a third sub-deep learning model to determine extra data for the user, classify the extra data as first extra result data if it is determined that the extra data does not exceed a previously set extra critical value, and classify the extra data as second extra result data if it is determined that the extra data exceeds the extra critical value, and the third sub-deep learning model is trained on multiple user data, multiple voice data, multiple tactile data, and multiple olfactory data, multiple extra data, multiple first extra result data, and multiple second extra result data.
[0060] According to various embodiments, the processor is configured to input the text data into the prediction model based on the following [Mathematical Formula 1] to generate a first quality loss function, a first cost loss function, a first efficiency loss function, and a first overall loss function for the text data; input the image data into the prediction model based on the following [Mathematical Formula 1] to generate a second quality loss function, a second cost loss function, a second efficiency loss function, and a second overall loss function for the image data; and input the extra data into the prediction model based on the following [Mathematical Formula 1] to generate a third quality loss function, a third cost loss function, a third efficiency loss function, and a third overall loss function for the extra data.
[0061] [Mathematical formula 1] JPEG2026079729000005.jpg41144 Here, L is the first overall loss function, the second overall loss function, and the third overall loss function, L quality L is the first quality loss function, the second quality loss function, and the third quality loss function. cost L is the first cost-loss function, the second cost-loss function, and the third cost-loss function, efficiencyis the first efficiency loss function, the second efficiency loss function, and the third efficiency loss function, where w1, w2, and w3 are the weighted values for the quality loss function, cost loss function, and efficiency loss function, respectively, L accuracy L is a value obtained by dividing the correct score for the domain. completeness is the completion score value for the domain, β1 and β2 are weighted values for calculating the quality loss function, and L token This is a calculated value relative to the price of the token, L cmputation is the sum of other processing costs excluding price, δ1 and δ2 are weighted values for calculating the aforementioned cost loss function, L time This is a ratio value to the actual processing time, L cmputation y1 is a ratio value to the number of tokens used, and y1 and y2 are weighted values for calculating the efficiency loss function.
[0062] According to various embodiments, the prediction model is learned based on a plurality of extraneous data, a plurality of first text result data, a plurality of second text result data, a plurality of first image result data, a plurality of second image result data, a plurality of first extraneous result data, and a plurality of second extraneous result data, in addition to a plurality of overall loss functions, a plurality of quality loss functions, a plurality of cost loss functions, and a plurality of efficiency loss functions.
[0063] According to various embodiments, the processor is configured to determine that the custom-made model is a user-customized model if the custom-made model is below a previously set first critical value, and to determine that some error has occurred in the custom-made model if the custom-made model is above the first critical value or below a second critical value set to be greater than the first critical value.
[0064] According to various embodiments, the processor determines that if the customized model is greater than or equal to the first critical value and the second critical value, the errors in the customized model are very serious, and only those customized models that are greater than or equal to the first critical value and the second critical value are to be retrained in the prediction model. The processor is also configured to send only those customized models that do not exceed the second critical value to the user's external electronic device via the communication interface.
[0065] Furthermore, according to various other embodiments, in a method for driving a server that recommends a user-customized model on a predictive model platform, the method is configured to receive at least one user data from the user's external electronic device via a communication interface, input the at least one user data into a first sub-deep learning model via a processor to determine text data for the user, input the at least one user data into a second sub-deep learning model via the processor to determine image data for the user, input at least one of the text data and / or image data into the predictive model via the processor to determine a custom-made model for the user, and send the custom-made model to the user's external electronic device via the communication interface, the predictive model is learned based on a plurality of text data and a plurality of image data, a plurality of custom-made models, first result data in which the plurality of custom-made models are judged to be accurate, second result data in which some errors are judged to have occurred in the plurality of custom-made models, and third result data in which errors are judged to have occurred in the plurality of custom-made models.
[0066] The terms “module” or “~part” as used in this document include units composed of hardware, software, or firmware, and can be used interchangeably with terms such as logic, logic block, component, or circuit. A “module” or “~part” can be a single component or the smallest unit or part thereof that performs one or more functions. A “module” or “~part” can be embodied mechanically or electronically and may include, for example, known or future ASIC (application-specific integrated circuit) chips, FPGAs (field-programmable gate arrays), or programmable logic devices that perform a certain operation and can be executed by processor 120. At least part of a device (e.g., a module or its function) or method (e.g., an operation) in various embodiments can be embodied in a set of instructions stored in a computer-readable storage medium (e.g., memory 130) in the form of a program module. When such instructions are executed by a processor (e.g., processor 120), the processor can perform the function corresponding to the instructions. Computer-readable recording media can include hard disks, floppy disks, magnetic media (e.g., magnetic tape), optical recording media (e.g., CD-ROMs, DVDs, magneto-optical media (e.g., floppy disks)), and internal memory. Instructions can include code created by a compiler or code that can be executed by an interpreter. Modules or program modules in various embodiments may include at least one of the aforementioned components, or some may be omitted, or may include other components. The operations performed by modules, program modules, or other components in various embodiments may be performed sequentially, in parallel, repeatedly, or heuristically, or at least some operations may be performed in other steps, or omitted, or other operations may be added.
[0067] Furthermore, the embodiments disclosed herein are presented for the purpose of explaining and understanding the disclosed and described content, and do not limit the scope of this disclosure. Therefore, the scope of this disclosure should be interpreted as including all modifications or various other embodiments based on the technical idea of this disclosure.
Claims
1. In a server that recommends user-customized models on a predictive model platform, Communication interface; Processor; including, The aforementioned processor, Through the aforementioned communication interface, the system receives at least one user data from the user's external electronic device. The data of at least one user is input into the first sub-deep learning model to determine the text data for the user. The aforementioned user data is input into the second sub-deep learning model to determine the image data for the user. Input at least one of the text data and / or image data into the prediction model to determine a customized model for the user, and The custom-made model is configured to be sent to the user's external electronic device via the aforementioned communication interface. The aforementioned prediction model, A server that learns based on multiple text data and multiple image data, multiple custom-made models, first result data in which the multiple custom-made models are judged to be accurate, second result data in which some errors are judged to have occurred in the multiple custom-made models, and third result data in which errors are judged to have occurred in the multiple custom-made models.
2. The aforementioned at least one user data is, The server according to claim 1, comprising: text data created by the user; translation data created by the user; summary data created by the user; dialogue data created by the user; audio data perceived or emitted by the user; tactile data perceived or emitted by the user; olfactory data perceived or emitted by the user; input image data confirmed by the user's eyes or entered by the user; analysis data of the user; and mathematical calculation data of the user.
3. The aforementioned processor, At least one of the aforementioned text data, translation data, summary data, and dialogue data is input to the first sub-deep learning model to determine the text data. If it is determined that the text data does not exceed the previously set text critical value, the text data is classified by the first text result data, and If it is determined that the text data exceeds the text critical value, the system is configured to classify the text data using the second text result data. The first sub-deep learning model described above is: The server according to claim 2, which is trained on multiple user data, multiple text data, multiple translation data, multiple summary data, and multiple dialogue data, and on multiple text data, multiple first text result data, and multiple second text result data.
4. The aforementioned processor, At least one of the input image data, the analysis data, and the at least one user data is input to the second sub-deep learning model to determine the image data. If it is determined that the image data does not exceed the previously set critical image value, the image data is classified by the first image result data, and If it is determined that the image data exceeds the image critical value, the system is configured to classify the image data using the second image result data. The second sub-deep learning model described above is: The server according to claim 3, which is trained on multiple user data, multiple analysis data, and multiple input image data, as well as multiple image data, multiple first image result data, and multiple second image result data.
5. The aforementioned processor, At least one of the aforementioned user data, voice data, tactile data, and olfactory data is input into the third sub-deep learning model to determine the extra data for the user. If it is determined that the aforementioned extra data does not exceed the previously set extra critical value, the aforementioned extra data is classified by the first extra result data, and If it is determined that the aforementioned extra data exceeds the aforementioned extra critical value, the system is set to classify the aforementioned extra data using the second extra result data. The third sub-deep learning model described above is: The server according to claim 4, which is trained on multiple user data, multiple voice data, multiple tactile data, and multiple olfactory data, and on multiple extra data, multiple first extra result data, and multiple second extra result data.
6. The aforementioned processor, The aforementioned text data is input into the prediction model based on the following [Mathematical Formula 1], and a first quality loss function, a first cost loss function, a first efficiency loss function, and a first overall loss function are generated for the text data, respectively. The aforementioned image data is input into the prediction model based on the following [Mathematical Formula 1] to generate a second quality loss function, a second cost loss function, a second efficiency loss function, and a second overall loss function for the aforementioned image data, and The aforementioned extra data is input into the prediction model based on the following [Mathematical Formula 1] to generate a third quality loss function, a third cost loss function, a third efficiency loss function, and a third overall loss function for the aforementioned extra data, respectively. [Mathematical formula 1] Here, L is the first overall loss function, the second overall loss function, and the third overall loss function, L quality is the first quality loss function, the second quality loss function, and the third quality loss function, L cost is the first cost loss function, the second cost loss function, and the third cost loss function, L efficiency is the first efficiency loss function, the second efficiency loss function, and the third efficiency loss function, w1, w2, w3 are the respective weighting values for the quality loss function, the cost loss function, and the efficiency loss function, L accuracy is the value obtained by dividing the correct score for the domain, L completeness is the completion score value for the domain, β 1 , β 2 is the weighting value for calculating the quality loss function, L token is the calculated value for the price comparison token, L overhead is the total value of other processing costs excluding the price, δ 1 , δ 2 is the weighting value for calculating the cost loss function, L time is the ratio value for the actual processing time, L cmputation is the ratio value for the number of tokens used, y1, y2 are the weighting values for calculating the efficiency loss function, The server according to claim 5.
7. The aforementioned prediction model, The server according to claim 6, which is further trained on multiple extra data, multiple first text result data, multiple second text result data, multiple first image result data, multiple second image result data, multiple first extra result data, and multiple second extra result data, and multiple overall loss functions, multiple quality loss functions, multiple cost loss functions, and multiple efficiency loss functions.
8. The aforementioned processor, If the custom-made model is less than or equal to the previously set first model critical value, then it is determined that the custom-made model is a user-customized model, and The server according to claim 6, wherein if the custom-made model is greater than or equal to the first critical value or less than or equal to a second critical value set to be greater than the first critical value, the server is configured to determine that a partial error has occurred in the custom-made model.
9. The aforementioned processor, If the customized model is above the first critical value and the second critical value, it is determined that the error in the customized model is very serious, and the predictive model is retrained only for customized models that are above the first critical value and the second critical value, and The server according to claim 8, configured to send the custom-made model to the user's external electronic device, but only for custom-made models that do not exceed the second critical value, via the communication interface.
10. In a method for driving a server that recommends user-customized models on a predictive model platform, The aforementioned method, Through the communication interface, receive at least one user data from the user's external electronic device, The data of at least one user is input to the first sub-deep learning model via the processor to determine the text data for the user. Through the processor, the data of at least one user is input to the second sub-deep learning model to determine the image data for the user. Through the processor, at least one of the text data and / or image data is input to the prediction model to determine a customized model for the user, and The custom-made model is configured to be sent to the user's external electronic device via the aforementioned communication interface. The aforementioned prediction model, The method according to claim 1, wherein the model is trained based on multiple text data and multiple image data, multiple custom-made models, first result data in which the multiple custom-made models are judged to be accurate, second result data in which some errors are judged to have occurred in the multiple custom-made models, and third result data in which errors are judged to have occurred in the multiple custom-made models.
11. The aforementioned at least one user data is, The method according to claim 10, comprising: text data created by the user; translation data created by the user; summary data created by the user; dialogue data created by the user; audio data perceived or emitted by the user; tactile data perceived or emitted by the user; olfactory data perceived or emitted by the user; input image data confirmed by the user's eyes or entered by the user; analysis data of the user; and mathematical calculation data of the user.
12. The aforementioned method, At least one of the aforementioned text data, translation data, summary data, and dialogue data is input into the first sub-deep learning model to determine the text data. If it is determined that the text data does not exceed the previously set text critical value, the text data is classified by the first text result data, and If it is determined that the text data exceeds the text critical value, the system is configured to classify the text data using the second text result data. The first sub-deep learning model described above is: The method according to claim 11, which is learned based on multiple user data, multiple text data, multiple translation data, multiple summary data, and multiple dialogue data, and multiple text data, multiple first text result data, and multiple second text result data.
13. The aforementioned method, At least one of the input image data, the analysis data, and the at least one user data is input to the second sub-deep learning model to determine the image data. If it is determined that the image data does not exceed the previously set critical image value, the image data is classified by the first image result data, and If it is determined that the image data exceeds the image critical value, the system is configured to classify the image data using the second image result data. The second sub-deep learning model described above is: The method according to claim 12, which is learned based on multiple user data, multiple analysis data, and multiple input image data, and multiple image data, multiple first image result data, and multiple second image result data.
14. The aforementioned method, At least one of the aforementioned user data, voice data, tactile data, and olfactory data is input into the third sub-deep learning model to determine the extra data for the user. If it is determined that the aforementioned extra data does not exceed the previously set extra critical value, the aforementioned extra data is classified by the first extra result data, and If it is determined that the aforementioned extra data exceeds the aforementioned extra critical value, the system is set to classify the aforementioned extra data using the second extra result data. The third sub-deep learning model described above is: The method according to claim 13, which is learned based on multiple user data, multiple voice data, multiple tactile data, and multiple olfactory data, and multiple extra data, multiple first extra result data, and multiple second extra result data.
15. The aforementioned method, The aforementioned text data is input into the prediction model based on the following [Mathematical Formula 1] to generate a first quality loss function, a first cost loss function, a first efficiency loss function, and a first overall loss function for the text data, respectively. The aforementioned image data is input into the prediction model based on the following [Mathematical Formula 1] to generate a second quality loss function, a second cost loss function, a second efficiency loss function, and a second overall loss function for the aforementioned image data, and The aforementioned extra data is input into the prediction model based on the following [Mathematical Formula 1], and the model is configured to generate a third quality loss function, a third cost loss function, a third efficiency loss function, and a third overall loss function for the aforementioned extra data. [Mathematical formula 1] Here, L is the first overall loss function, the second overall loss function, and the third overall loss function, L quality L is the first quality loss function, the second quality loss function, and the third quality loss function. cost L is the first cost-loss function, the second cost-loss function, and the third cost-loss function. efficiency L is the first efficiency loss function, the second efficiency loss function, and the third efficiency loss function, where w1, w2, and w3 are the weighted values for the quality loss function, cost loss function, and efficiency loss function, respectively. accuracy L is a value obtained by dividing the correct score for the domain. completeness This is a score representing the level of completion for the domain, and β 1 ,β 2 L is a weighted value for calculating the aforementioned quality loss function. token This is a calculated value relative to the price of the token, L overhead This is the sum of other processing costs excluding price, δ 1 , δ 2 L is a weighted value for calculating the aforementioned cost loss function. time This is a ratio value to the actual processing time, L cmputation The method according to claim 14, wherein y is a ratio value to the number of tokens used, and y1 and y2 are weighted values for calculating the efficiency loss function.
16. The aforementioned prediction model, The method according to claim 15, wherein the model is learned based on a plurality of extra data, a plurality of first text result data, a plurality of second text result data, a plurality of first image result data, a plurality of second image result data, a plurality of first extra result data, and a plurality of second extra result data, in addition to a plurality of overall loss functions, a plurality of quality loss functions, a plurality of cost loss functions, and a plurality of efficiency loss functions.
17. The aforementioned method, If the custom-made model is less than or equal to the previously set first model critical value, then it is determined that the custom-made model is a user-customized model, and The method according to claim 15, wherein if the custom-made model is greater than or equal to the first critical value or less than or equal to a second critical value set to be greater than the first critical value, it is determined that a partial error has occurred in the custom-made model.
18. The aforementioned method, If the customized model is above the first critical value and the second critical value, it is determined that the error in the customized model is very serious, and the predictive model is retrained only for customized models that are above the first critical value and the second critical value, and The method according to claim 17, wherein the custom-made model is configured to be sent to the user's external electronic device only if the custom-made model does not exceed the second critical value, via the communication interface.