Electronic device and operation method therefor

WO2026160938A1PCT designated stage Publication Date: 2026-07-30SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2026-01-26
Publication Date
2026-07-30

Smart Images

  • Figure KR2026001518_30072026_PF_FP_ABST
    Figure KR2026001518_30072026_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device according to one embodiment disclosed herein identifies one or more objects through each of a first artificial intelligence model and a second artificial intelligence model on the basis of a content image input to the electronic device, and transmits the content image and information corresponding to the identified objects to a server on the basis that the identified objects do not correspond to each other.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device and method of operation thereof

[0001] The present disclosure relates to an electronic device and a method of operating the same, a server and a method of operating the same, a system including an electronic device and a server, a computer-readable recording medium storing a computer program for operating the method of operating the electronic device, and a computer-readable recording medium storing a computer program for operating the method of operating the server.

[0002] In the fields of image processing and computer vision, Artificial Intelligence (AI) has achieved a level of performance improvement that was previously impossible. However, AI-based image processing algorithms had a limitation in that they required a high amount of computation. Recently, along with the lightweighting of these algorithms, the performance improvement and optimization of hardware for their computation have enabled the realization of on-device methods that perform AI-based image processing within the device. The on-device approach refers to a method in which AI-based algorithms are executed directly on the device itself, such as smartphones, tablets, and IoT (Internet of Things) devices, rather than on a cloud server.

[0003] There may be a difference between the dataset used for training at the time of development of an on-device model and the dataset actually provided when the on-device model is deployed and run directly on the device. Accordingly, research is being conducted to address performance degradation caused by differences in datasets even when the on-device model is run directly on the device.

[0004] According to one embodiment of the present disclosure, an electronic device may be provided.

[0005] An electronic device according to one embodiment of the present disclosure may include a communication interface comprising a communication circuit, a memory storing a plurality of instructions, and at least one processor comprising a processing circuitry and operably coupled to the memory.

[0006] An electronic device according to one embodiment of the present disclosure can identify one or more objects through each of a first artificial intelligence model and a second artificial intelligence model based on a content image input to the electronic device, by having a plurality of instructions executed individually or collectively by at least one processor.

[0007] An electronic device according to one embodiment of the present disclosure can transmit a content image, information corresponding to one or more objects identified through a first artificial intelligence model, and information corresponding to one or more objects identified through a second artificial intelligence model to a server through a communication interface, based on the fact that one or more objects identified through a first artificial intelligence model and one or more objects identified through a second artificial intelligence model do not correspond to each other by executing a plurality of instructions individually or collectively by at least one processor.

[0008] An electronic device according to one embodiment of the present disclosure can receive from a server, via a communication interface, information corresponding to an update of at least one of a first artificial intelligence model or a second artificial intelligence model, which is obtained using a training image generated based on a content image, information corresponding to one or more objects identified through a first artificial intelligence model, and information corresponding to one or more objects identified through a second artificial intelligence model, by having a plurality of instructions executed individually or collectively by at least one processor.

[0009] According to one embodiment of the present disclosure, a method of operating an electronic device may be provided.

[0010] A method of operation of an electronic device according to one embodiment of the present disclosure may include the step of identifying one or more objects through each of a first artificial intelligence model and a second artificial intelligence model based on a content image input to the electronic device.

[0011] A method of operation of an electronic device according to one embodiment of the present disclosure may include the step of transmitting to a server a content image, information corresponding to one or more objects identified through a first artificial intelligence model, and information corresponding to one or more objects identified through a second artificial intelligence model, based on the fact that one or more objects identified through a first artificial intelligence model and one or more objects identified through a second artificial intelligence model do not correspond to each other.

[0012] A method of operation of an electronic device according to one embodiment of the present disclosure may include the step of receiving from a server information corresponding to an update of at least one of a first artificial intelligence model or a second artificial intelligence model, which is obtained using a learning image based on a content image, information corresponding to one or more objects identified through a first artificial intelligence model, and information corresponding to one or more objects identified through a second artificial intelligence model.

[0013] According to one embodiment of the present disclosure, a computer-readable recording medium may be provided that records a program for executing any one of the methods of the electronic device described above and below.

[0014] The aspects, features, and advantages of the above and other embodiments of the present disclosure will become more apparent from the following detailed description together with the accompanying drawings, where reference numerals denote structural elements.

[0015] FIG. 1 is a conceptual diagram illustrating an exemplary operation of receiving update information regarding a model stored in an electronic device from a server according to one embodiment of the present disclosure.

[0016] FIG. 2 is a flowchart illustrating an exemplary operation of an electronic device according to one embodiment of the present disclosure.

[0017] FIG. 3 is a flowchart illustrating an exemplary method of operating a server according to one embodiment of the present disclosure.

[0018] FIG. 4 is a flowchart illustrating an exemplary method of operating a system according to one embodiment of the present disclosure.

[0019] FIG. 5a is a block diagram showing an exemplary configuration of an electronic device according to one embodiment of the present disclosure.

[0020] FIG. 5b is a block diagram showing an exemplary configuration of a server according to one embodiment of the present disclosure.

[0021] FIG. 5c is a block diagram showing an exemplary configuration of a system according to one embodiment of the present disclosure.

[0022] FIG. 6a is a flowchart illustrating an exemplary operation of transmitting information regarding an object identified through content images and models of an electronic device according to one embodiment of the present disclosure to a server.

[0023] FIG. 6b is a diagram illustrating an exemplary operation of transmitting information regarding an object identified through content images and models of an electronic device according to one embodiment of the present disclosure to a server.

[0024] FIG. 7a is a flowchart illustrating an exemplary operation of transmitting information regarding an object identified through content images and models of an electronic device according to one embodiment of the present disclosure to a server.

[0025] FIG. 7b is a diagram illustrating an exemplary operation of transmitting information regarding an object identified through content images and models of an electronic device according to one embodiment of the present disclosure to a server.

[0026] FIG. 8a is a flowchart illustrating an exemplary operation of generating a training image of a server according to one embodiment of the present disclosure.

[0027] FIG. 8b is a flowchart illustrating an exemplary operation of generating a training image of a server according to one embodiment of the present disclosure.

[0028] FIG. 8c is a diagram illustrating an exemplary operation of generating a training image of a server according to one embodiment of the present disclosure.

[0029] FIG. 9a is a flowchart illustrating an exemplary operation of determining the identification difficulty of each of the content image and the learning image of a server according to one embodiment of the present disclosure, and storing the learning image.

[0030] FIG. 9b is a diagram illustrating an exemplary operation of determining the identification difficulty of each of the content image and the learning image of a server according to one embodiment of the present disclosure, and storing the learning image.

[0031] FIG. 9c is a diagram illustrating an exemplary operation of determining the identification difficulty of each of the content image and the learning image of a server according to one embodiment of the present disclosure, and storing the learning image.

[0032] FIG. 10a is a flowchart illustrating an exemplary operation of learning a model of a server according to one embodiment of the present disclosure.

[0033] FIG. 10b is a flowchart illustrating an exemplary operation of learning a model of a server according to one embodiment of the present disclosure.

[0034] FIG. 10c is a diagram illustrating an exemplary operation of learning a model of a server according to one embodiment of the present disclosure.

[0035] FIG. 11a is a flowchart illustrating an exemplary operation of distributing a model to a plurality of electronic devices of a server according to one embodiment of the present disclosure.

[0036] FIG. 11b is a diagram illustrating an exemplary operation of distributing a model to a plurality of electronic devices of a server according to one embodiment of the present disclosure.

[0037] FIG. 12a is a flowchart illustrating an exemplary operation of learning a model of an electronic device according to one embodiment of the present disclosure.

[0038] FIG. 12b is a diagram illustrating an exemplary operation of learning a model of an electronic device according to one embodiment of the present disclosure.

[0039] FIG. 13 is a block diagram showing an exemplary configuration of an electronic device according to one embodiment of the present disclosure.

[0040] In the present disclosure, the expression “at least one of a, b, or c” may refer to “a”, “b”, “c”, “a and b”, “a and c”, “b and c”, “a, b, and c all”, or variations thereof.

[0041] Embodiments of the present disclosure are described below in detail with reference to the attached drawings so that those skilled in the art can easily implement them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein.

[0042] The terms used in this disclosure are described in their current, general form considering the functions mentioned herein; however, they may refer to various other terms depending on the intent of those skilled in the art, case law, the emergence of new technologies, etc. Accordingly, the terms used in this disclosure should not be interpreted solely by their names, but should be interpreted based on the meaning of the terms and the overall content of this disclosure.

[0043] Furthermore, the terms used in this disclosure are used merely to describe specific embodiments and are not intended to limit this disclosure.

[0044] Throughout the entire disclosure, when a part is described as being "connected" to another part, this includes not only cases where they are "directly connected," but also cases where they are "electrically connected" with other elements interposed between them.

[0045] The terms “above” and similar designations used in the present disclosure, particularly in the claims, may indicate both singular and plural forms. Furthermore, unless there is a description explicitly specifying the order of the steps describing the method according to the present disclosure, the described steps may be performed in a suitable order. The present disclosure is not limited by the order in which the described steps are described.

[0046] Phrases such as "in one embodiment" appearing in various places in this specification do not necessarily refer to the same embodiment.

[0047] Embodiments of the present disclosure may be represented by functional block configurations and various processing steps. Some or all of these functional blocks may be implemented by various numbers of hardware and / or software configurations that execute specific functions. For example, the functional blocks of the present disclosure may be implemented by one or more microprocessors or by circuit configurations for a specific (e.g., designated, predetermined, or (pre)determined) function. Additionally, for example, the functional blocks of the present disclosure may be implemented in various programming or scripting languages. The functional blocks may be implemented as algorithms executed on one or more processors. Furthermore, the present disclosure may employ prior art for electronic configuration, signal processing, and / or data processing, etc. Terms such as “mechanism,” “element,” “means,” and “configuration” may be used broadly and are not limited to mechanical and physical configurations.

[0048] The connecting lines or connecting members between the components depicted in the drawings are merely illustrative of functional connections and / or physical or circuit connections. In the actual device, connections between components may be represented by various alternative or additional functional, physical, or circuit connections.

[0049] Additionally, terms such as "...part," "module," etc., as described in the specification refer to a unit that processes at least one function or operation, and this may be implemented in hardware or software, or as a combination of hardware and software.

[0050] In the present disclosure, the “processor” may include various processing circuits and / or a plurality of processors. For example, the term “processor” as used herein, including in the claims, may include at least one processor and various processing circuits. In the at least one processor, one or more processors may be configured to perform the various functions described herein in a distributed manner, individually and / or collectively. As used herein, the “processor,” “at least one processor,” and “one or more processors” may be configured to perform various functions. However, these terms cover, without limitation, situations where one processor performs some of the functions and other processor(s) perform other parts of the functions, and situations where a single processor can perform all functions. Additionally, the at least one processor may include a combination of processors performing various functions of the disclosed functions in a distributed manner. The at least one processor may execute program instructions to achieve or perform various functions.

[0051] In the present disclosure, artificial intelligence technology may include machine learning (deep learning) technology that utilizes algorithms for self-classification / learning of the characteristics of input data, and elemental technologies that utilize machine learning algorithms to mimic functions such as cognition and judgment of the human brain. The elemental technologies may include, for example, at least one of linguistic understanding technology that recognizes human language / characters, visual understanding technology that recognizes objects like human vision, inference / prediction technology that judges information to logically infer and predict, knowledge representation technology that processes human experience information into knowledge data, and motion control technology that controls autonomous driving of vehicles and the movement of robots. Linguistic understanding is a technology that recognizes, applies, and processes human language / characters, and may include natural language processing, machine translation, dialogue systems, question answering, speech recognition / synthesis, etc. Visual understanding is a technology that recognizes and processes objects like human vision, and may include object recognition, object tracking, image search, person recognition, scene understanding, spatial understanding, image enhancement, etc. Inference and prediction is a technology that judges information to logically infer and predict, and may include knowledge / probability-based inference, optimization prediction, preference-based planning, recommendation, etc. Knowledge representation is a technology that automatically processes human experiential information into knowledge data and may include knowledge construction (data generation / classification) and knowledge management (data utilization).

[0052] A predefined rule of action or an artificial intelligence model may be characterized as being created through learning. Here, being created through learning means, for example, that a basic artificial intelligence model is trained using a number of training data by a learning algorithm, thereby creating a predefined rule of action or an artificial intelligence model configured to perform a desired characteristic (or objective). Such learning may be performed on the device itself where the artificial intelligence according to the present disclosure is executed, or it may be performed through a separate server and / or system. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but are not limited to the examples described above.

[0053] An artificial intelligence model may include multiple neural network layers. Each of the multiple neural network layers has multiple weight values ​​and can perform neural network operations through operations between the results of previous layers and the multiple weights. The multiple weights possessed by the multiple neural network layers can be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights can be updated so that the loss value or cost value obtained from the artificial intelligence model during the learning process is reduced or minimized. The artificial neural network may include a Deep Neural Network (DNN), such as a Convolutional Neural Network (CNN), Deep Neural Network (DNN), Recurrent Neural Network (RNN), Restricted Boltzmann Machine (RBM), Deep Belief Network (DBN), Bidirectional Recurrent Deep Neural Network (BRDNN), or Deep Q-Networks, but is not limited to the examples mentioned above.

[0054] The present disclosure will be described in more detail below with reference to the attached drawings.

[0055] FIG. 1 is a diagram illustrating an exemplary operation of receiving update information regarding a model stored in an electronic device (1000) from a server (2000) within a system (100) according to one embodiment of the present disclosure.

[0056] Referring to FIG. 1, a system (100) according to one embodiment of the present disclosure may include at least one electronic device (1000) and a server (2000). Although FIG. 1 shows only one electronic device (1000) for convenience of explanation, the system (100) may include a plurality of electronic devices (1000). That is, the server (2000) may transmit and receive data with a plurality of electronic devices (1000).

[0057] In a system (100) according to one embodiment of the present disclosure, if it is identified that an electronic device (1000) has obtained result data of an incorrect answer through at least one model stored in the electronic device (1000), the electronic device (1000) may transmit information related to the result data of an incorrect answer to a server (2000). In a system (100) according to one embodiment of the present disclosure, the server (2000) may generate a learning image (2) corresponding to a content image (1) based on information received from the electronic device (1000). The server (2000) may transmit to the electronic device (1000) information (3) corresponding to an update of at least one model stored in the electronic device (1000), which is obtained using the learning image (2).

[0058] An electronic device (1000) according to one embodiment of the present disclosure may be implemented as an electronic device (1000) of various types and forms including a display. The electronic device (1000) may include devices capable of displaying through a display, such as a smart TV, smartphone, tablet PC, PDA (personal digital assistant), laptop PC, glasses-type display, and head-mounted display (HMD), but is not limited thereto. For example, the electronic device (1000) may be implemented as an electronic device (1000) of various types and forms capable of wired / wireless connection with a display. For example, the electronic device (1000) may include devices capable of displaying through wired / wireless connection with a display, such as a set-top box or desktop PC, but is not limited thereto.

[0059] In one embodiment of the present disclosure, an electronic device (1000) may acquire a content image (1). At this time, the content image (1) may be a still image at a specific point in time among time-series images of content provided through the electronic device (1000). At this time, "content" may refer to various things that can be executed by the electronic device (1000) and provided to a user as video or sound, such as movies, dramas, animations, applications, novels, comic books, advertisements, or web pages. For example, when the electronic device (1000) receives an input regarding an object detection request, it may acquire a still image at the point in time when the input is received among the time-series images of the provided content.

[0060] The electronic device (1000) may store at least one model that recognizes one or more objects in an input image. That is, the electronic device (1000) may be equipped with at least one model that identifies objects. In the present disclosure, 'object' may refer to a specific object within an image (or video) and may be classified by class. For example, a person, animal, object, natural object, building, etc. within the image (or video) may be the subject of the object. In one embodiment of the present disclosure, the object may include a specific area of ​​a specific object. For example, the object may include a human face.

[0061] In one embodiment of the present disclosure, an electronic device (1000) may store a first artificial intelligence model (10) and a second artificial intelligence model (20). In the present disclosure, the first artificial intelligence model (10) may also be referred to as the first model, and the second artificial intelligence model (20) may also be referred to as the second model. The first artificial intelligence model (10) and the second artificial intelligence model (20) may be artificial intelligence models that identify one or more objects in an input image. The first artificial intelligence model (10) and the second artificial intelligence model (20) may be different types of artificial intelligence models.

[0062] In one embodiment of the present disclosure, an electronic device (1000) can identify at least one object from a content image (1) using a first artificial intelligence model (10). An electronic device (1000) can identify at least one object from a content image (1) using a second artificial intelligence model (20). The electronic device (1000) can compare the object identification result of the first artificial intelligence model (10) and the object identification result of the second artificial intelligence model (20) from the content image (1), and if the two results are identified as different, the electronic device (1000) can determine that the first artificial intelligence model (10) and / or the second artificial intelligence model (20) detected incorrect result data in identifying an object within the content image (1). If the electronic device (1000) determines that incorrect result data has been detected from the first artificial intelligence model (10) and / or the second artificial intelligence model (20), it can transmit information related to the incorrect result data to a server (2000). For example, information related to the result data of an incorrect answer may include a content image (1) that is the target from which the result data of an incorrect answer was detected, information (11) corresponding to at least one object identified through the first artificial intelligence model (10) from the content image (1), and information (21) corresponding to at least one object identified through the second artificial intelligence model (20) from the content image (1).

[0063] In one embodiment of the present disclosure, a server (2000) may receive information regarding at least one object identified from a content image (1) through at least one artificial intelligence model (e.g., a first artificial intelligence model (10), a second artificial intelligence model (20)) stored in an electronic device (1000) from an electronic device (1000).

[0064] In one embodiment of the present disclosure, at least one model for identifying one or more objects in an input image may be stored in the server (2000). That is, the server (2000) may be equipped with at least one model for identifying objects. In one embodiment of the present disclosure, a third artificial intelligence model (30) may be stored in the server (2000). In the present disclosure, the third artificial intelligence model (30) may also be referred to as the third model. The third artificial intelligence model (30) may be an artificial intelligence model that identifies one or more objects in an input image. The third artificial intelligence model (30) may be an artificial intelligence model of a different type from the first artificial intelligence model (10) and the second artificial intelligence model (20). The server (2000) may be a device with higher computing performance than the electronic device (1000) so as to be able to perform more operations quickly than the electronic device (1000). Accordingly, the model installed on the server (2000) (e.g., the third artificial intelligence model (30)) can have higher performance than the models installed on the electronic device (1000) (e.g., the first artificial intelligence model (10) and the second artificial intelligence model (20)).

[0065] In one embodiment of the present disclosure, a server (2000) can identify an object from a content image (1) using a third artificial intelligence model (30). The server (2000) can compare the object identification result of the third artificial intelligence model (30) with the object identification results of each of the first and second artificial intelligence models (10, 20) from the content image (1). The server (2000) can extract an incorrect answer area (31) from the content image (1) that corresponds to an object identified differently by the first and second artificial intelligence models (10, 20) mounted on the electronic device (1000) and the third artificial intelligence model (30) mounted on the server (2000).

[0066] In one embodiment of the present disclosure, a server (2000) may generate and store a learning image (2) based on a content image (1) received from an electronic device (1000) and an incorrect answer area (31) extracted from the content image (1). The learning image (2) may be an image based on the content image (1) and an incorrect answer area (31) corresponding to an object identified differently between the electronic device (1000) and the server (2000). At this time, as the server (2000) generates the learning image (2) based on the extracted incorrect answer area (31), the incorrect answer area (31) may be concretized and expressed in the generated learning image (2). The operation and method of generating the learning image (2) will be described in more detail later with reference to FIGS. 8a to 8c.

[0067] In one embodiment of the present disclosure, a server (2000) can learn at least one model (e.g., a first artificial intelligence model (10) and a second artificial intelligence model (20)) stored in an electronic device (1000) using a stored training image (2). An operation and method for performing model learning will be described in more detail later with reference to FIGS. 10a to 10c.

[0068] In one embodiment of the present disclosure, a server (2000) may receive information corresponding to an update of at least one of the models stored in the electronic device (1000), which is derived by learning at least one model (e.g., a first artificial intelligence model (10) and a second artificial intelligence model (20)) stored in the electronic device (1000) using a generated training image (2). For example, the information corresponding to an update of at least one of the models stored in the electronic device (1000) may include at least one of a model updated based on the training image (2) based on the first artificial intelligence model (10) or a model updated based on the training image (2) based on the second artificial intelligence model (20). An operation and method for distributing (or updating) the model after training the model will be described in more detail below with reference to FIGS. 11a and FIGS. 11b.

[0069] According to one embodiment of the present disclosure, the server (2000) performs model learning based on a content image (1), which is the target of the detection of incorrect answer result data in the electronic device (1000), thereby improving the object identification performance of artificial intelligence models (e.g., first and second artificial intelligence models (10, 20)) mounted on the electronic device (1000).

[0070] According to one embodiment of the present disclosure, a server (2000) can generate a learning image (2) based on a content image (1) actually provided by an electronic device (1000) and use the generated learning image (2) to learn artificial intelligence models (e.g., first and second artificial intelligence models (10, 20)) installed on the electronic device (1000). Accordingly, the server (2000) can learn the artificial intelligence models (10, 20) without licensing issues such as copyright by using the learning image (2) generated within the server (2000). Meanwhile, the dataset used for learning at the time of development of the on-device model may differ from the dataset actually provided at the time when the on-device model is actually distributed to the electronic device (1000) and directly operated within the electronic device (1000). However, in learning artificial intelligence models (10, 20), the server (2000) according to one embodiment of the present disclosure may improve model performance in an actual environment by using a learning image (2) generated based on a content image (1) actually provided by an electronic device (1000).

[0071] According to one embodiment of the present disclosure, as the server (2000) generates a training image (2) in which the incorrect answer area (31) is materialized, model training can be performed to improve object identification performance in the incorrect answer area (31).

[0072] FIG. 2 is a flowchart illustrating an exemplary method of operating an electronic device (1000) according to one embodiment of the present disclosure. Hereinafter, in describing a method of operating an electronic device (1000) according to one embodiment of the present disclosure, reference may be made to FIG. 1 and FIG. 2 together.

[0073] In step S210 of FIG. 2, the electronic device (1000) can identify one or more objects through each of the first artificial intelligence model (10) and the second artificial intelligence model (20) based on the input content image (1).

[0074] In one embodiment of the present disclosure, an electronic device (1000) may store a first artificial intelligence model (10) and a second artificial intelligence model (20). The first artificial intelligence model (10) and the second artificial intelligence model (20) may be artificial intelligence models that identify one or more objects in an input image. The first artificial intelligence model (10) and the second artificial intelligence model (20) may be different types of artificial intelligence models.

[0075] In the present disclosure, identifying objects may mean determining where objects are located in a given image (object localization) and determining which category each object belongs to (object classification). In one embodiment of the present disclosure, artificial intelligence models for identifying objects may undergo three steps, for example, informative region selection, feature extraction from each candidate region, and class classification of the object candidate regions by applying a classifier to the extracted features. Depending on the detection method, localization performance may be improved through subsequent post-processing, such as bounding box regression.

[0076] In one embodiment of the present disclosure, the first artificial intelligence model (10) may be a small model, and the second artificial intelligence model (20) may be a middle model. In the present disclosure, 'small model' refers to a small artificial intelligence model and may refer to a model having a relatively small number of parameters and a relatively low computational complexity. In the present disclosure, 'middle model' refers to an artificial intelligence model of medium size and may refer to a model having more parameters and a higher computational complexity than the small model. The small model may perform faster computational processing than the middle model, but may provide lower performance than the middle model. The middle model may provide higher performance than the small model, but may have slower computational processing than the small model.

[0077] In one embodiment of the present disclosure, the first artificial intelligence model (10) may be a first type of small model, and the second artificial intelligence model (20) may be a second type of small model. That is, the first artificial intelligence model (10) and the second artificial intelligence model (20) may both be small models, but may be different types of artificial intelligence models.

[0078] In one embodiment of the present disclosure, an electronic device (1000) may obtain information (11) corresponding to one or more objects identified from a content image (1) through a first artificial intelligence model (10). The first artificial intelligence model (10) may receive a content image (1), identify one or more objects within the content image (1), and output information (11) corresponding to the identified one or more objects. For example, the first artificial intelligence model (10) may output class information of an object and location information of an object as information (11) corresponding to one or more objects recognized from the content image (1).

[0079] The first artificial intelligence model (10) can perform an algorithm to detect one or more objects within an image based on an input image. The first artificial intelligence model (10) may be a pre-trained artificial intelligence model capable of identifying one or more objects within an image according to the input image and outputting information regarding the identified one or more objects.

[0080] In one embodiment of the present disclosure, an electronic device (1000) may obtain information (21) corresponding to one or more objects identified from a content image (1) through a second artificial intelligence model (20). The second artificial intelligence model (20) may receive a content image (1) as input, identify one or more objects within the content image (1), and output information (21) corresponding to the identified one or more objects. For example, the second artificial intelligence model (20) may output class information of an object and location information of an object as information (21) corresponding to one or more objects identified from the content image (1).

[0081] The second artificial intelligence model (20) can perform an algorithm to detect one or more objects within an image based on an input image. The second artificial intelligence model (20) may be a pre-trained artificial intelligence model capable of identifying one or more objects within an image according to the input image and outputting information regarding one or more recognized objects.

[0082] In step S220 of FIG. 2, if one or more objects identified through the first artificial intelligence model (10) and one or more objects identified through the second artificial intelligence model (20) do not correspond to each other, the electronic device (1000) can transmit a content image (1), information (11) corresponding to one or more objects identified through the first artificial intelligence model (10), and information (21) corresponding to one or more objects identified through the second artificial intelligence model (20) to the server (2000). Based on the fact that one or more objects identified through the first artificial intelligence model (10) and one or more objects identified through the second artificial intelligence model (20) do not correspond to each other, the electronic device (1000) can transmit a content image (1), information (11) corresponding to one or more objects identified through the first artificial intelligence model (10), and information (21) corresponding to one or more objects identified through the second artificial intelligence model (20) to the server (2000).

[0083] In one embodiment of the present disclosure, an electronic device (1000) can compare information (11) corresponding to one or more objects identified through a first artificial intelligence model (10) with information (21) corresponding to one or more objects identified through a second artificial intelligence model (20). For example, regarding a specific object identified through the second artificial intelligence model (20), the electronic device (1000) can identify that one or more objects identified through the first artificial intelligence model (10) and one or more objects identified through the second artificial intelligence model (20) do not correspond to each other if the object is not identified through the first artificial intelligence model (10). For example, regarding a specific object identified through the second artificial intelligence model (20), if the electronic device (1000) is also identified as an object through the first artificial intelligence model (10) but is identified as a different class, it can identify that one or more objects identified through the first artificial intelligence model (10) and one or more objects identified through the second artificial intelligence model (20) do not correspond to each other.

[0084] In step S230 of FIG. 2, the electronic device (1000) may receive from the server (2000) information (3) corresponding to an update of at least one of the first artificial intelligence model (10) or the second artificial intelligence model (20), which is obtained using a training image generated based on a content image (1), information (11) corresponding to one or more objects identified through the first artificial intelligence model (10), and information (21) corresponding to one or more objects identified through the second artificial intelligence model (20).

[0085] In one embodiment of the present disclosure, an electronic device (1000) may receive information (3) corresponding to an update of at least one of a first artificial intelligence model (10) or a second artificial intelligence model (20) from a server (2000). The first artificial intelligence model (10) and the second artificial intelligence model (20) mounted on the electronic device (1000) may be learned using the server (2000).

[0086] In one embodiment of the present disclosure, information (3) corresponding to an update of an artificial intelligence model may include information regarding whether an update of the artificial intelligence model is necessary and information regarding updated parameters if it is identified that an update is necessary. In one embodiment of the present disclosure, information (3) corresponding to an update of an artificial intelligence model may include information regarding whether an update of the artificial intelligence model is necessary and, if it is identified that an update is necessary, the artificial intelligence model itself trained on the server (2000). In this case, the trained artificial intelligence model itself may be provided in a file format.

[0087] In one embodiment of the present disclosure, the electronic device (1000) can update a corresponding artificial intelligence model based on information (3) corresponding to an update of at least one of a first artificial intelligence model (10) and a second artificial intelligence model (20) received from a server (2000).

[0088] FIG. 3 is a flowchart illustrating an exemplary method of operating a server (2000) according to one embodiment of the present disclosure. Hereinafter, in describing the method of operating a server (2000) according to one embodiment of the present disclosure, the description will be made with reference to FIG. 1 and FIG. 3 together.

[0089] In step S310 of FIG. 3, the server (2000) can receive information corresponding to one or more objects identified from the content image (1) through the content image (1) and one or more artificial intelligence models (e.g., a first artificial intelligence model (10) and / or a second artificial intelligence model (20)) stored in the at least one electronic device (1000) from at least one electronic device (1000).

[0090] In one embodiment of the present disclosure, when it is determined that incorrect answer result data is detected from an artificial intelligence model stored in an electronic device (1000), the server (2000) may receive from at least one electronic device (1000) a content image (1) to which the incorrect answer result data is detected and the incorrect answer result data. At this time, the incorrect answer result data may include information regarding one or more objects identified from the content image (1) through the artificial intelligence model stored in the electronic device (1000) (for example, information (11) corresponding to one or more objects identified through the first artificial intelligence model (10) and / or information (21) corresponding to one or more objects identified through the second artificial intelligence model (20)).

[0091] In step S320 of FIG. 3, the server (2000) can identify one or more objects from the content image (1) through an artificial intelligence model (or a third artificial intelligence model (30)) stored in the server (2000) and extract an incorrect answer area (31) corresponding to an object identified differently from each other in at least one electronic device (1000) and the server (2000) from the content image (1).

[0092] In one embodiment of the present disclosure, the artificial intelligence model (30) stored in the server (2000) may be a large model. In the present disclosure, 'large model' may mean a large artificial intelligence model, a model having a larger number of parameters and a more complex architecture than a small model and a middle model. The server (2000) may be a device with higher computing performance than the electronic device (1000) so as to be able to perform more computations quickly than the electronic device (1000). Accordingly, the artificial intelligence model (or the third artificial intelligence model (30)) stored in the server (2000) may provide higher performance than one or more artificial intelligence models stored in the electronic device (1000). For example, the object identification performance of an artificial intelligence model (e.g., a third artificial intelligence model (30)) stored in a server (2000) may be higher than the object identification performance of one or more artificial intelligence models (e.g., first and second artificial intelligence models (10, 20)) stored in an electronic device (1000).

[0093] In one embodiment of the present disclosure, one or more objects can be identified from a content image (1) using an artificial intelligence model (30) stored in a server (2000). If the object result identified through the artificial intelligence model (30) stored in the server (2000) and the object result identified through one or more artificial intelligence models stored in the electronic device (1000) are different from each other, the object result identified through the artificial intelligence model (or the third artificial intelligence model (30)) stored in the server (2000) may be considered as the response (or correct answer). The server (2000) may extract an area corresponding to an object identified differently from each other in at least one electronic device (1000) and the server (2000) in the content image (1) as an incorrect answer area (31).

[0094] In step S330 of Fig. 3, the server (2000) can generate and store a learning image (2) based on the content image (1) and the incorrect answer area (31).

[0095] In one embodiment of the present disclosure, the server (2000) can generate a base image based on the remaining area excluding the incorrect answer area (31) of the content image (1), and can generate a learning image (2) by refining the base image based on the incorrect answer area (31) of the content image (1). The learning image (2) can correspond to an image obtained by refining the base image based on features extracted from the remaining area excluding the incorrect answer area (31) of the content image (1) based on features extracted from the incorrect answer area (31). Through this, the server (2000) can generate a learning image (2) in which the incorrect answer area (31) of the content image (1) is depicted in detail.

[0096] In step S340 of FIG. 3, the server (2000) can learn a model corresponding to one or more artificial intelligence models stored in the electronic device (1000) using the stored training image (2).

[0097] In one embodiment of the present disclosure, a first artificial intelligence model (10) and a second artificial intelligence model (20) mounted on an electronic device (1000) may be trained using a server (2000). The server (2000) may store a model corresponding to the first artificial intelligence model (10) and a model corresponding to the second artificial intelligence model (20). In the present disclosure, the inclusion of a model corresponding to the artificial intelligence model mounted on the electronic device (1000) in the server may mean that the artificial intelligence model stored in the server (2000) and the artificial intelligence model stored in the electronic device (1000) share the same architecture and the same weights (or have the same architecture and the same model parameters). For example, the server (2000) may store a model identical to the first artificial intelligence model (10) mounted (or stored) on the electronic device (1000). For example, the server (2000) may have a model identical to the second artificial intelligence model (20) mounted (or stored) on the electronic device (1000). The server (2000) can learn the model corresponding to the first artificial intelligence model (10) and the model corresponding to the second artificial intelligence model (20) using the generated training image (2).

[0098] FIG. 4 is a flowchart illustrating an exemplary method of operation of a system (100) according to one embodiment of the present disclosure. Hereinafter, in describing the method of operation of an electronic device (1000) according to one embodiment of the present disclosure, the description will be made with reference to FIG. 1 and FIG. 4 together. However, regarding steps S410 to S470 shown in FIG. 4, the description in FIG. 2 and FIG. 3 applies identically, so the description of such steps will not be repeated.

[0099] In step S410 of FIG. 4, the electronic device (1000) can identify one or more objects from an input content image (1) using each of the first artificial intelligence model (10) and the second artificial intelligence model (20). In step S420 of FIG. 4, the electronic device (1000) can compare the object identification result of the first artificial intelligence model (10) with the object identification result of the second artificial intelligence model (20). In step S420 of FIG. 4, if it is determined that the object identification result of the first artificial intelligence model (10) and the object identification result of the second artificial intelligence model (20) are different from each other, the electronic device (1000) can proceed to step S430 to transmit the content image (1), information (11) corresponding to one or more objects identified through the first artificial intelligence model (10), and information (21) corresponding to one or more objects identified through the second artificial intelligence model (20) to the server (2000).

[0100] In step S440 of FIG. 4, the server (2000) can identify one or more objects from the content image (1) through the third artificial intelligence model (30) and extract an incorrect answer area (31). The incorrect answer area (31) can be extracted as an area corresponding to an object identified differently from each other by at least one electronic device (1000) and the server (2000) from the content image (1).

[0101] In step S450 of FIG. 4, the server (2000) can generate a training image (2) based on the content image (1) and the incorrect answer area (31). In step S460 of FIG. 4, the server (2000) can use the training image (2) to train a model corresponding to each of the first artificial intelligence model (10) and the second artificial intelligence model (20).

[0102] In step S470 of FIG. 4, the server (2000) can transmit information corresponding to an update of at least one of the first artificial intelligence model (10) or the second artificial intelligence model (20) to the electronic device (1000).

[0103] FIG. 5a is a block diagram illustrating an exemplary configuration of an electronic device (1000) according to one embodiment of the present disclosure.

[0104] Referring to FIG. 5a, an electronic device (1000) according to one embodiment of the present disclosure may include a communication interface (110) (e.g., including a communication circuit), a processor (120) (e.g., including a processing circuit), and a memory (130).

[0105] The communication interface (110) may include various communication circuits and can perform data communication with the server (2000) under the control of the processor (120).

[0106] The communication interface (110) may include a communication circuit. The communication interface (110) may include a communication circuit capable of performing data communication between an electronic device (1000) and other devices using at least one of a data communication method including, for example, wired LAN, wireless LAN, Wi-Fi, Bluetooth, ZigBee, WFD (Wi-Fi Direct), infrared communication (IrDA, infrared Data Association), BLE (Bluetooth Low Energy), NFC (Near Field Communication), Wibro (Wireless Broadband Internet), WiMAX (World Interoperability for Microwave Access), SWAP (Shared Wireless Access Protocol), WiGig (Wireless Gigabit Alliances, WiGig) and RF communication.

[0107] The electronic device (1000) can transmit information regarding the result of object identification of a content image performed within the electronic device (1000) to the server (2000) using the communication interface (110). The electronic device (1000) can receive information from the server (2000) corresponding to an update of at least one artificial intelligence model mounted on the electronic device (1000) using the communication interface (110).

[0108] The memory (130) can store a program for processing and controlling the processor (120), and can store data that is input to or output from the electronic device (1000). Additionally, the memory (130) can store data necessary for the operation of the electronic device (1000).

[0109] The memory (130) may include at least one type of storage medium among, for example, a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), RAM (Random Access Memory), SRAM (Static Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), PROM (Programmable Read-Only Memory), magnetic memory, a magnetic disk, or an optical disk.

[0110] The processor (120) may include various processing circuits and controls the operations of the electronic device (1000). For example, the processor (120) may perform the functions of the electronic device (1000) described in the present disclosure by executing one or more instructions stored in memory (130).

[0111] In an embodiment of the present disclosure, the processor (120) may store one or more instructions in an internally provided memory (130) and control the operation of an electronic device (1000) to be performed by executing one or more instructions stored in the internally provided memory (130). That is, the processor (120) may perform a predetermined operation by executing at least one instruction or program stored in an internal memory or memory (130) provided within the processor (120).

[0112] The processor (120) may include various processing circuits and / or multiple processors. For example, the term "processor" as used herein, including in the claims, may include at least one processor and various processing circuits. In at least one processor, one or more processors may be configured to perform the various functions described herein in a distributed manner, individually and / or collectively. As used herein, "processor," "at least one processor," and "one or more processors" may be configured to perform various functions. However, these terms cover, for example but without limitation, situations where one processor performs some of the functions and other processor(s) perform other parts of the functions, and situations where a single processor can perform all functions. Additionally, at least one processor may include a combination of processors performing various functions of the disclosed functions in a distributed manner. At least one processor may execute program instructions to achieve or perform various functions.

[0113] By executing one or more instructions individually or collectively by at least one processor (120), an electronic device (1000) according to one embodiment of the present disclosure can identify one or more objects through each of a first artificial intelligence model (132) and a second artificial intelligence model (133) based on an input content image. Each model may include various circuits and / or executable program instructions. If one or more objects identified through the first artificial intelligence model (132) and one or more objects identified through the second artificial intelligence model (133) do not correspond to each other, the electronic device (1000) according to one embodiment of the present disclosure can transmit the content image, information corresponding to one or more objects identified through the first artificial intelligence model (132), and information corresponding to one or more objects identified through the second artificial intelligence model (133) to a server (2000) via a communication interface (110). An electronic device (1000) according to one embodiment of the present disclosure may receive from a server (2000) information corresponding to an update of at least one of a first artificial intelligence model or a second artificial intelligence model, obtained using a learning image based on a content image, information corresponding to one or more objects identified through a first artificial intelligence model (132), and information corresponding to one or more objects identified through a second artificial intelligence model (133), through a communication interface (110).

[0114] An electronic device (1000) according to one embodiment of the present disclosure may receive an input corresponding to object identification for each of a plurality of content images. The electronic device (1000) according to one embodiment of the present disclosure may execute a first artificial intelligence model (132) and a second artificial intelligence model (133) respectively upon receiving an input corresponding to object identification when the resolution of each of the plurality of content images is less than a threshold value, and when the resolution of each of the plurality of content images is greater than or equal to a threshold value, execute the first artificial intelligence model (132) upon receiving an input corresponding to object identification and execute the second artificial intelligence model (133) in response to a preset frequency.

[0115] An electronic device (1000) according to one embodiment of the present disclosure may receive an input corresponding to object identification for each of a plurality of content images. An electronic device (1000) according to one embodiment of the present disclosure may alternately execute a first artificial intelligence model (132) and a second artificial intelligence model (133) upon receiving an input corresponding to object identification. An electronic device (1000) according to one embodiment of the present disclosure may execute both the first artificial intelligence model (132) and the second artificial intelligence model (133) when the frequency of execution of the first artificial intelligence model (132) and the second artificial intelligence model (133) corresponds to a preset frequency.

[0116] An electronic device (1000) according to one embodiment of the present disclosure may store information corresponding to a content image and one or more objects identified through the second artificial intelligence model (133) when one or more objects identified through the first artificial intelligence model (132) and one or more objects identified through the second artificial intelligence model (133) do not correspond to each other. An electronic device (1000) according to one embodiment of the present disclosure may learn the first artificial intelligence model (131) by inputting information corresponding to one or more objects identified through the second artificial intelligence model (133) as a response to object identification of the content image.

[0117] According to one embodiment of the present disclosure, information corresponding to an update of at least one of the first artificial intelligence model (132) and the second artificial intelligence model (133) may include at least one of a model updated based on information corresponding to a learning image and one or more objects identified from the learning image through the server (2000) based on the first artificial intelligence model (131), or a model updated based on information corresponding to a learning image and one or more objects identified from the learning image through the server (2000) based on the second artificial intelligence model (132).

[0118] According to one embodiment of the present disclosure, the learning image may be based on an incorrect answer area and a content image corresponding to an object that is differently identified between an electronic device (1000) and a server (2000).

[0119] According to one embodiment of the present disclosure, a learning image may correspond to an image materialized based on a second feature prompt corresponding to an incorrect answer area, from a base image based on a first feature prompt corresponding to the remaining area excluding the incorrect answer area among the content images.

[0120] According to one embodiment of the present disclosure, a content image may include a first content image having a first identification difficulty based on the fact that information corresponding to one or more objects identified from a (first) content image through a first artificial intelligence model (132) and information corresponding to one or more objects identified from a (first) content image through an artificial intelligence model stored in a server (2000) are identical, and a second content image having a second identification difficulty based on the fact that information corresponding to one or more objects identified from a (second) content image through a second artificial intelligence model (133) and information corresponding to one or more objects identified from a (second) content image through an artificial intelligence model stored in a server (2000) are identical. According to one embodiment of the present disclosure, a learning image may include a first learning image corresponding to a first content image having a first identification difficulty and a second learning image corresponding to a second content image having a second identification difficulty.

[0121] An electronic device (1000) according to one embodiment of the present disclosure may receive from a server (2000) either a first type artificial intelligence model or a second type artificial intelligence model trained based on each of a first type dataset and a second type dataset, wherein the ratio between a first training image and a second training image is different from each other based on a first artificial intelligence model (132). An electronic device (1000) according to one embodiment of the present disclosure may transmit the test result of either the first type artificial intelligence model or the second type artificial intelligence model to a server.

[0122] An electronic device (1000) according to one embodiment of the present disclosure may receive from a server information corresponding to a final model determined based on the test results of either the first type artificial intelligence model or the second type artificial intelligence model among the first artificial intelligence model (132), the first type artificial intelligence model, and the second type artificial intelligence model.

[0123] An electronic device (1000) according to one embodiment of the present disclosure can update a first artificial intelligence model (132) based on information corresponding to a received final model.

[0124] FIG. 5b is a block diagram illustrating an exemplary configuration of a server (2000) according to one embodiment of the present disclosure.

[0125] Referring to FIG. 5b, a server (2000) according to one embodiment of the present disclosure may include a communication interface (210) (e.g., including a communication circuit), a processor (220) (e.g., including a processing circuit), and a memory (230).

[0126] The communication interface (210) may include various communication circuits and may perform data communication with at least one electronic device (1000) under the control of the processor (220).

[0127] The communication interface (210) may include a communication circuit. The communication interface (210) may include a communication circuit capable of performing data communication between a server (2000) and other devices using at least one of a data communication method including, for example, wired LAN, wireless LAN, Wi-Fi, Bluetooth, ZigBee, WFD (Wi-Fi Direct), infrared communication (IrDA, infrared Data Association), BLE (Bluetooth Low Energy), NFC (Near Field Communication), Wibro (Wireless Broadband Internet), WiMAX (World Interoperability for Microwave Access), SWAP (Shared Wireless Access Protocol), WiGig (Wireless Gigabit Alliances), and RF communication.

[0128] The server (2000) can receive information regarding the object identification result of a content image performed inside the electronic device (1000) from the electronic device (1000) using the communication interface (210). The server (2000) can transmit information corresponding to an update of at least one artificial intelligence model mounted on the electronic device (1000) to the electronic device (1000) using the communication interface (210).

[0129] The memory (230) can store a program for processing and controlling the processor (220), and can store data that is input to or output from the server (2000). Additionally, the memory (230) can store data necessary for the operation of the server (2000).

[0130] The memory (230) may include at least one type of storage medium among, for example, a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), RAM (Random Access Memory), SRAM (Static Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), PROM (Programmable Read-Only Memory), magnetic memory, a magnetic disk, and an optical disk.

[0131] The processor (220) may include various processing circuits and may control the operations of the server (2000). For example, the processor (220) may perform the functions of the server (2000) described in the present disclosure by executing one or more instructions stored in memory (230).

[0132] In an embodiment of the present disclosure, the processor (220) may store one or more instructions in an internally provided memory (230) and control the execution of operations of the server (2000) by executing one or more instructions stored in the internally provided memory (230). That is, the processor (220) may perform a predetermined operation by executing at least one instruction or program stored in an internal memory or memory (230) provided within the processor (220).

[0133] The processor (220) may include various processing circuits and / or multiple processors. For example, the term "processor" as used herein, including in the claims, may include at least one processor and various processing circuits. In at least one processor, one or more processors may be configured to perform the various functions described herein individually and / or collectively in a distributed manner. As used herein, "processor," "at least one processor," and "one or more processors" may be configured to perform various functions. However, these terms cover, for example but without limitation, situations where one processor performs some of the functions and other processor(s) perform other parts of the functions, and situations where a single processor can perform all functions. Additionally, at least one processor may include a combination of processors performing various functions of the disclosed functions in a distributed manner. At least one processor may execute program instructions to achieve or perform various functions.

[0134] By having one or more instructions executed individually or collectively by at least one processor (220), a server (2000) according to one embodiment of the present disclosure may receive from at least one electronic device (1000) information regarding one or more objects identified from a content image through one or more artificial intelligence models stored in at least one electronic device (1000). The server (2000) may identify one or more objects from a content image through an artificial intelligence model stored in the server (2000) (each of which may include various circuits and / or executable program instructions) and extract an incorrect answer area from the content image corresponding to one or more objects identified differently in at least one electronic device (1000) and the server (2000). Based on the content image and the incorrect answer area, the server (2000) may generate and store a training image. The server (2000) may learn one or more models corresponding to one or more artificial intelligence models stored in the electronic device (1000) using the stored training image.

[0135] By executing one or more instructions individually or collectively by at least one processor (220), a server (2000) according to one embodiment of the present disclosure can generate a base image based on features extracted from the remaining area excluding the incorrect answer area of ​​the content image. The server (2000) can generate a learning image by concretizing the base image based on features extracted from the incorrect answer area of ​​the content image.

[0136] By executing one or more instructions individually or collectively by at least one processor (220), a server (2000) according to one embodiment of the present disclosure may acquire or generate a first feature prompt corresponding to the remaining area of ​​a content image. Based on the first feature prompt, the server (2000) may generate a base image. The server (2000) may acquire or generate a second feature prompt corresponding to the incorrect area of ​​the content image. Based on the second feature prompt, the server (2000) may generate a training image from the base image.

[0137] By executing one or more instructions individually or collectively by at least one processor (220), a server (2000) according to one embodiment of the present disclosure can identify one or more objects from a training image through an artificial intelligence model stored in the server (2000) and store information corresponding to one or more objects identified from the training image as a response (or correct answer) to the object identification of the training image.

[0138] By executing one or more instructions individually or collectively by at least one processor (220), a server (2000) according to one embodiment of the present disclosure may determine the identification difficulty of a content image as a first identification difficulty if the information corresponding to one or more objects identified from a content image through a first artificial intelligence model (132) and the information corresponding to one or more objects identified from a content image through an artificial intelligence model stored in the server (2000) are identical. The server (2000) may determine the identification difficulty of a content image as a second identification difficulty if the information corresponding to one or more objects identified from a content image through a second artificial intelligence model and the information corresponding to one or more objects identified from a content image through an artificial intelligence model stored in the server (2000) are identical. The server (2000) may determine the identification difficulty of a learning image as a first identification difficulty if the content image corresponding to the learning image has a first identification difficulty, and may determine the identification difficulty of a learning image as a second identification difficulty if the content image corresponding to the learning image has a second identification difficulty.

[0139] By executing one or more instructions individually or collectively by at least one processor (220), a server (2000) according to one embodiment of the present disclosure can learn a model corresponding to a first artificial intelligence model based on each of a first type dataset and a second type dataset having different ratios between a training image having a first identification difficulty and a training image having a second identification difficulty.

[0140] By having one or more instructions executed individually or collectively by at least one processor (220), a server (2000) according to one embodiment of the present disclosure may store a first type artificial intelligence model learned based on a first type dataset and a second type artificial intelligence model learned based on a second type dataset. The server (2000) may distribute the first type artificial intelligence model to some of at least one electronic device (1000) and distribute the second type artificial intelligence model to other parts of at least one electronic device. The server (2000) may receive information regarding the test results of the first type artificial intelligence model from some of the at least one electronic device (1000). The server (2000) may receive information regarding the test results of the second type artificial intelligence model from other parts of at least one electronic device (1000). The server (2000) may receive information regarding the test results of the first artificial intelligence model from the remainder of at least one electronic device (1000), excluding some and other parts. The server (2000) can determine a final model among the first artificial intelligence model, the first artificial intelligence model, and the second artificial intelligence model based on information regarding the test results of the first type artificial intelligence model, information regarding the test results of the second type artificial intelligence model, and information regarding the test results of the first artificial intelligence model. If the determined final model is the first type artificial intelligence model or the second type artificial intelligence model, the server (2000) can distribute the determined first type artificial intelligence model or the second type artificial intelligence model to at least one electronic device (1000).

[0141] By executing one or more instructions individually or collectively by at least one processor (220), a server (2000) according to one embodiment of the present disclosure can learn a model corresponding to a second artificial intelligence model (133) based on each of a third type dataset and a fourth type dataset having different ratios between a training image having a first identification difficulty and a training image having a second identification difficulty.

[0142] According to one embodiment of the present disclosure, a content image received from at least one electronic device (1000) may correspond to an image in which one or more objects identified through a first artificial intelligence model (132) and one or more objects identified through a second artificial intelligence model (133) are different from each other.

[0143] In one embodiment of the present disclosure, the memory (230) may include an object identification module (231). The 'module' included in the memory (230) may refer to a unit that processes a function or operation performed by the processor (220), and may be implemented as software such as instructions, algorithms, data structures, or program code.

[0144] The object identification module (231) may include appropriate logic, circuits, interfaces, and / or code that can be operated to identify (or detect) one or more objects in an input image using one or more artificial intelligence models. In one embodiment of the present disclosure, the object identification module (231) stored in a server (2000) may include a third artificial intelligence model (232). For example, the third artificial intelligence model (232) may correspond to a large model.

[0145] Meanwhile, although not shown, the server (2000) can perform model learning for each of the first artificial intelligence model (132) and the second artificial intelligence model (133) mounted on the electronic device (1000), and the memory (230) may include a model corresponding to the first artificial intelligence model (132) and a model corresponding to the second artificial intelligence model (133).

[0146] Meanwhile, although not shown, the memory (230) may further include a feature prompt extraction module. The feature prompt extraction module may include appropriate logic, circuits, interfaces, and / or code that can be operated to generate an image-generating prompt from the input image based on features extracted from the input image.

[0147] Meanwhile, although not shown, the memory (230) may further include an image generation module. The image generation module may include appropriate logic, circuits, interfaces, and / or code that can be operated to generate a new image based on a feature prompt using one or more neural networks.

[0148] FIG. 5c is a block diagram illustrating an exemplary configuration of a system (3000) according to one embodiment of the present disclosure.

[0149] Referring to FIG. 5c, in one embodiment of the present disclosure, the system (3000) may include an electronic device (1000) and a server (2000) connected to a communication network. The electronic device (1000) according to one embodiment of the present disclosure may include a communication interface (110), a processor (120), and a memory (130). The server (2000) according to one embodiment of the present disclosure may include a communication interface (210), a processor (220), and a memory (230). However, regarding the configurations illustrated in FIG. 5c, the descriptions in FIG. 5a and FIG. 5b apply equally, so the descriptions regarding such configurations will not be repeated.

[0150] FIG. 6a is a flowchart illustrating an exemplary operation of transmitting information regarding an object recognized through content images and models of an electronic device (1000) according to one embodiment of the present disclosure to a server (2000). FIG. 6b is a diagram illustrating an exemplary operation of transmitting information regarding one or more objects recognized through content images and models of an electronic device (1000) according to one embodiment of the present disclosure to a server (2000).

[0151] In step S610 of FIG. 6a, the electronic device (1000) can receive an input corresponding to object identification for each of the plurality of content images. For example, the electronic device (1000) can provide a user interface that can input a request for object identification for each of the plurality of content images, and can receive a user input requesting object identification through the user interface.

[0152] Referring together with FIG. 6b, an electronic device (1000) according to one embodiment of the present disclosure can restore data of a content image (601) received in a compressed format to the original image format before compression through a decoder (602). The decoder (602) can interpret encoded input data and decode image data (or video data) included in the input data using a decoding method suitable for the input data.

[0153] In one embodiment of the present disclosure, the electronic device (1000) may provide an object identification service (603). For example, the electronic device (1000) may provide a sports broadcast video and, at the same time, identify a player appearing in the sports broadcast video and provide a user interface containing information about the player by overlaying it on the sports broadcast video. For example, the electronic device (1000) may provide a multimedia content video and, at the same time, identify an entertainer such as an actor appearing in the multimedia content video and provide a user interface containing information about the entertainer by overlaying it on the multimedia content video. For example, the electronic device (1000) may obtain information search results regarding a person appearing in the video through a search server or a cloud server.

[0154] In one embodiment of the present disclosure, an electronic device (1000) may receive an input (604) corresponding to object identification based on the activation of an object identification service (603). Although FIG. 6b illustrates the electronic device (1000) receiving a single content image (601), in practice, it may receive consecutive frames as content images are provided. Accordingly, the electronic device (1000) may receive an input (604) corresponding to object identification for each of a plurality of content images (601) corresponding to consecutive frames.

[0155] In one embodiment of the present disclosure, the electronic device (1000) can obtain resolution information for each of a plurality of content images.

[0156] In one embodiment of the present disclosure, an electronic device (1000) may obtain resolution information (605) of a content image (601) (hereinafter also referred to as content resolution information (605)) from a decoder (602). Considering that the response time of the model varies according to the resolution information (605) of the content image (601), the electronic device (1000) may determine the frequency of the middle model (or, execution frequency, 610) based on the resolution information (605) of the content image (601).

[0157] In step S620 of FIG. 6a, the electronic device (1000) can execute the first artificial intelligence model and the second artificial intelligence model, respectively, upon receiving an input corresponding to object identification, if the resolution of the plurality of content images is less than a threshold value. In step S620 of FIG. 6a, the electronic device (1000) can execute the first artificial intelligence model upon receiving an input corresponding to object identification and execute the second artificial intelligence model in response to a preset frequency, if the resolution of the plurality of content images is greater than or equal to a threshold value. The electronic device (1000) can execute the first artificial intelligence model and the second artificial intelligence model, respectively, upon receiving an input corresponding to object identification, based on the fact that the resolution of the plurality of content images is less than a threshold value. The electronic device (1000) can execute the first artificial intelligence model upon receiving an input corresponding to object identification and execute the second artificial intelligence model in response to a predetermined frequency, based on the fact that the resolution of the plurality of content images is greater than or equal to a threshold value.

[0158] Referring together to FIG. 6a and FIG. 6b, in one embodiment of the present disclosure, an object identification module (606) within an electronic device (1000) may include a first artificial intelligence model and a second artificial intelligence model, and FIG. 6b illustrates, as an example, that the first artificial intelligence model is a small model (607) and the second artificial intelligence model is a middle model (608).

[0159] In one embodiment of the present disclosure, the electronic device (1000) may perform object identification using a small model (607) upon receiving an input corresponding to each object identification. The electronic device (1000) may perform object identification using the small model (607) whenever an input corresponding to each object identification is received. For example, the small model (607) may be a model actually used when object identification is to be performed in providing the object identification service (603) of the electronic device (1000).

[0160] In one embodiment of the present disclosure, the electronic device (1000) can identify (609) whether the resolution is within a range that can be processed within a preset response time based on content resolution information (605) obtained from the decoder (602).

[0161] For example, it is assumed that the response time constraint of the object identification module (606) is 150ms. If the content resolution is a first resolution (e.g., 640x320 resolution), the time required for object identification of the small model (607) may be 20ms, and the time required for object identification of the middle model (608) may be 100ms. In this case, the electronic device (1000) can identify that the total time required for object identification of the small model (607) and the middle model (608) does not exceed the response time constraint of the object identification module (606). Accordingly, the electronic device (1000) can execute the middle model (608) upon receiving an input corresponding to each object identification, just like the small model (607), based on the fact that the content resolution is within a range that can be processed within a preset response time. The electronic device (1000) can execute the middle model (608) whenever an input corresponding to each object identification is received, just like the small model (607), based on the fact that the content resolution is within a range that can be processed within a preset response time.

[0162] When the content resolution is a second resolution (e.g., 1280x720 resolution) that is higher than the first resolution, the time required for object identification of the small model (607) may be 80ms, and the time required for object identification of the middle model (608) may be 400ms. In this case, the electronic device (1000) may identify that the total time required for object identification of the small model (607) and the middle model (608) exceeds the response time constraint of the object identification module (606). Accordingly, the electronic device (1000) may call the middle model (608) at a predetermined (e.g., specified, predetermined, (pre) determined, preset) frequency based on the fact that the content resolution exceeds the range that can be processed within the preset response time.

[0163] In one embodiment of the present disclosure, an electronic device (1000) may receive a preset frequency (610) from a server (2000). For example, the execution frequency (610) may be provided in time units, and the electronic device (1000) may execute a middle model (608) at a preset (e.g., specified, preset, (pre)determined, preset) time interval. For example, the server (2000) may be configured to execute the middle model (608) at a first time interval (e.g., 1 minute) when the dataset is collected in a small amount (or in an amount less than a threshold) based on the amount of collected dataset, to perform object identification for 1 frame per minute. For example, the server (2000) can be configured to perform object identification for 1 frame per 5 minutes by running the middle model (608) every second time (e.g., 5 minutes) which is longer than the first time, based on the amount of collected dataset (or, if the amount collected is greater than a threshold amount).

[0164] When both the small model (607) and the middle model (608) are executed in the object identification module (606), the electronic device (1000) can compare (611) the object identification result of the content image (601) in the small model (607) with the object identification result of the content image (601) in the middle model (608). If the electronic device (1000) identifies that the object identification result in the small model (607) and the object identification result in the middle model (608) are different for the content image (601), it can transmit to the server (2000) the content image (601), information (612) corresponding to the object identified through the small model (607), and information (613) corresponding to the object identified through the middle model (608). If the electronic device (1000) does not correspond to the information (612) corresponding to the object identified through the small model (607) and the information (613) corresponding to the object identified through the middle model (608) for the content image (601), the content image (601), the information (612) corresponding to the object identified through the small model (607), and the information (613) corresponding to the object identified through the middle model (608) can be transmitted to the server (2000). For example, each of the information (612) corresponding to the object identified through the small model (607) and the information (613) corresponding to the object identified through the middle model (608) may include object class information and object location information.

[0165] If the electronic device (1000) identifies that the object identification result in the small model (607) and the object identification result in the middle model (608) for the content image (601) are identical to each other, it is considered that the object identification of the content image (601) in both the small model (607) and the middle model (608) was performed correctly, and thus the electronic device (1000) may not transmit separate information to the server (2000). If the information (612) corresponding to the object identified through the small model (607) and the information (613) corresponding to the object identified through the middle model (608) for the content image (601) correspond to each other, the electronic device (1000) may not transmit separate information to the server (2000).

[0166] FIG. 7a is a flowchart illustrating an exemplary operation of transmitting information regarding an object recognized through content images and models of an electronic device (1000) according to one embodiment of the present disclosure to a server (2000). FIG. 7b is a diagram illustrating an exemplary operation of transmitting information regarding an object recognized through content images and models of an electronic device (1000) according to one embodiment of the present disclosure to a server (2000).

[0167] In step S710 of FIG. 7a, the electronic device (1000) can receive an input corresponding to object identification for each of the plurality of content images.

[0168] Referring together with FIG. 7b, in one embodiment of the present disclosure, an electronic device (1000) may receive an input (703) corresponding to object identification based on the activation of an object identification service (702). FIG. 7b illustrates the electronic device (1000) receiving a single content image (701), but substantially, it may receive consecutive frames as content images are provided. Accordingly, the electronic device (1000) may receive an input (703) corresponding to object identification for each of a plurality of content images (701) corresponding to consecutive frames.

[0169] In step S720 of FIG. 7a, the electronic device (1000) can alternately execute the first artificial intelligence model and the second artificial intelligence model upon receiving an input corresponding to object identification. In step S730 of FIG. 7a, the electronic device (1000) can execute both the first artificial intelligence model and the second artificial intelligence model if the frequency of execution of the first artificial intelligence model and the second artificial intelligence model corresponds to a preset frequency. The electronic device (1000) can execute both the first artificial intelligence model and the second artificial intelligence model based on the fact that the frequency of execution of the first artificial intelligence model and the second artificial intelligence model corresponds to a predetermined frequency.

[0170] Referring together with FIG. 7b, in one embodiment of the present disclosure, an object identification module (704) within an electronic device (1000) may include a first artificial intelligence model and a second artificial intelligence model, and FIG. 7b illustrates, as an example, that both the first artificial intelligence model and the second artificial intelligence model are small models, but are different types of models. The first artificial intelligence model may be referred to as a first type of small model (705), and the second artificial intelligence model may be referred to as a second type of small model (706). For example, the first type of small model (705) may be an EfficientDet object detection model, and the second type of small model (706) may be a Damo-YOLO object detection model.

[0171] In one embodiment of the present disclosure, the electronic device (1000) may alternately execute a first type of small model (705) and a second type of small model (706) upon receiving an input corresponding to each object identification. The electronic device (1000) may alternately execute a first type of small model (705) and a second type of small model (706) whenever an input corresponding to each object identification is received. For example, both the first type of small model (705) and the second type of small model (706) may be models actually used when object identification is to be performed in providing the object identification service (702) of the electronic device (1000).

[0172] In one embodiment of the present disclosure, an electronic device (1000) may receive a predetermined (e.g., specified, predetermined, (pre)determined, preset) frequency (707) from a server (2000). For example, the predetermined (e.g., specified, predetermined, (pre)determined, preset) frequency (707) may be provided in time units, and the electronic device (1000) may execute both a first type of small model (705) and a second type of small model (706) at predetermined (e.g., specified, predetermined, (pre)determined, preset) times. For example, the server (2000) may be configured to execute both a first type of small model (705) and a second type of small model (706) at a first time (e.g., 1 minute) when the amount of the collected dataset is small (or, less than a threshold amount). For example, the server (2000) can be configured to run both the first type of small model (705) and the second type of small model (706) at a second time (e.g., 5 minutes) longer than the first time, based on the amount of collected dataset.

[0173] When both the first type of small model (705) and the second type of small model (706) are executed in the object identification module (704), the electronic device (1000) can compare (708) the object identification result of the content image (701) in the first type of small model (705) with the object identification result of the content image (701) in the second type of small model (706). If the electronic device (1000) identifies that the object identification result in the first type of small model (705) and the object identification result in the second type of small model (706) are different for the content image (701), the electronic device (1000) can determine that the content image (701) is an image for which object identification is difficult. Accordingly, the electronic device (1000) can transmit the content image (701), information (709) corresponding to an object identified through a first type small model, and information (710) corresponding to an object identified through a second type small model to the server (2000). If the information (709) corresponding to an object identified through a first type small model (705) and the information (710) corresponding to an object identified through a second type small model (706) do not correspond to each other with respect to the content image (701), the electronic device (1000) can transmit the content image (701), information (709) corresponding to an object identified through a first type small model, and information (710) corresponding to an object identified through a second type small model to the server (2000). For example, each of the information (709) corresponding to an object identified through the first type small model (705) and the information (710) corresponding to an object identified through the second type small model (706) may include object class information and object location information.

[0174] If the electronic device (1000) identifies that the object identification result in the first type of small model (705) and the object identification result in the second type of small model (706) for the content image (701) are identical to each other, it is considered that the object identification of the content image (701) in both the first type of small model (705) and the second type of small model (706) was performed correctly, and thus the electronic device (1000) may not transmit information regarding the content image to the server (2000) performing model learning. If the information (709) corresponding to the object identified through the first type of small model (705) for the content image (701) and the information (710) corresponding to the object identified through the second type of small model (706) correspond to each other, the electronic device (1000) may not transmit information regarding the content image to the server (2000) performing model learning.

[0175] Hereinafter, with reference to FIGS. 8a to 8c, the operation of a server (2000) that generates a learning image based on a content image received from an electronic device (1000) will be described in more detail.

[0176] FIG. 8a is a flowchart illustrating an exemplary operation of generating a training image of a server (2000) according to one embodiment of the present disclosure. FIG. 8b is a flowchart illustrating an exemplary operation of generating a training image of a server (2000) according to one embodiment of the present disclosure. FIG. 8c is a diagram illustrating an exemplary operation of generating a training image of a server (2000) according to one embodiment of the present disclosure.

[0177] Step S810 of FIG. 8a is a step that embodies Step S330 of FIG. 3. The operation of Step S820 illustrated in FIG. 8a can be performed after the operation of Step S340 illustrated in FIG. 3 has been performed. The operation of Step S320 illustrated in FIG. 3 can be performed after the operation of Step S830 illustrated in FIG. 8a has been performed.

[0178] In step S810 of FIG. 8a, the server (2000) can generate a base image based on features extracted from the remaining area excluding the incorrect answer area of ​​the content image.

[0179] In step S820 of FIG. 8a, the server (2000) can generate a training image by refining the base image based on features extracted from the incorrect answer area among the content images.

[0180] In one embodiment of the present disclosure, the server (2000) can extract an incorrect answer area from a content image that corresponds to an object different from an object identified by an artificial intelligence model mounted on the electronic device (1000) among one or more objects identified by an artificial intelligence model mounted on the server (2000). That is, the incorrect answer area may represent an area containing an object that is difficult to identify by one or more artificial intelligence models mounted on the electronic device (1000).

[0181] In one embodiment of the present disclosure, the training image may be based on an incorrect answer area and a content image corresponding to an object that is differently identified between an electronic device and a server. In one embodiment of the present disclosure, the training image may correspond to an image materialized based on features extracted from an incorrect answer area, from a base image based on features extracted from the remaining area excluding the incorrect answer area among the content images.

[0182] According to one embodiment of the present disclosure, the server (2000) can generate a training image from a content image received from an electronic device (1000) by distinguishing and describing the incorrect answer area from the remaining area excluding the incorrect answer area among the content image, thereby generating a training image in which the incorrect answer area is expressed in more detail than when the entire content image is described at once. The server (2000) can train an artificial intelligence model mounted on the electronic device (1000) using the training image in which the incorrect answer area is expressed in detail, and thereby, the identification performance of the incorrect answer area that could not be identified previously can be improved.

[0183] Steps S811 and S812 of FIG. 8b are steps that embody step S810 of FIG. 8a. Steps S821 and S822 of FIG. 8b are steps that embody step S820 of FIG. 8a. The operation of step S811 illustrated in FIG. 8a can be performed after the operation of step S320 illustrated in FIG. 3 has been performed. The operation of step S340 illustrated in FIG. 3 can be performed after the operation of step S822 illustrated in FIG. 8a has been performed.

[0184] In step S811 of FIG. 8b, the server (2000) may generate a first feature prompt corresponding to the remaining area of ​​the content image. In step S812 of FIG. 8b, the server (2000) may generate a base image based on the first feature prompt. In the present disclosure, the first feature prompt may also be referred to as a base image feature prompt.

[0185] In the present disclosure, a 'prompt' may be used as input information required for a generative model to perform a task. The prompt may include natural language text. The natural language text may include various information such as a task indicating the task the generative model is to perform, and a context, intent, constraints, and example as components available when the generative model performs the task. In one embodiment of the present disclosure, an electronic device (1000) may process natural language text using a Natural Language Processing (NLP) model. In the present disclosure, the prompt may be replaced with an input, command, directive, input phrase, starting sentence, task query, trigger sentence, etc.

[0186] In one embodiment of the present disclosure, the prompt may include a multimedia prompt that integrates various types of media elements, including text, images, voice, video, music, animation, etc. The multimedia prompt may be a combination of different types of media elements in the same situation.

[0187] In one embodiment of the present disclosure, a first feature prompt (or, base image feature prompt) is created based on features extracted (or, depicted) from an input image, and a server (2000) can generate a base image based on the first feature prompt (or, base image feature prompt).

[0188] Referring together to FIG. 8b and FIG. 8c, a server (2000) according to one embodiment of the present disclosure can generate a base image feature prompt (804) through a prompt generation module (810). The server (2000) can generate a base image feature prompt (804) from a content image (801) with an incorrect answer area (802) masked through the prompt generation module (810). In generating the base image feature prompt (804), the server (2000) can extract features regarding the remaining area (803) excluding the incorrect answer area (802) by using the content image (801) with the incorrect answer area (802) masked. For example, the prompt generation module (810) can generate a base image feature prompt (804) in which the overall features of the image are described by extracting the overall features of the image from an input image. For example, the prompt generation module (810) can generate a base image feature prompt (804) that describes the composition, background, foreground, person, action, lighting, tone, texture, atmosphere, etc. within the image from the input image.

[0189] The server (2000) can generate a base image feature prompt (804) with the content "A stone breakwater extends long toward the sea against the backdrop of an open beach. Small pebbles and water remain on the ground, showing traces of recent rain or splashing seawater. Gentle waves are crashing on the sea, and clouds are hanging over the cloudy sky. A seagull is seen flying in the distance, and the overall atmosphere is calm and quiet. The horizon where the sea and sky meet is clearly visible, and the blue of the sea and the gray of the sky are in natural harmony. In the image, a man is standing on the right and looking to the left. The man is wearing an ivory knit sweater, gray slacks, and white sneakers. The man has his hands in his pockets."

[0190] In the base image feature prompt (804) presented as an example above, the first and second sentences describe the features regarding the background and foreground of the input image, and the third and fourth sentences describe the features regarding the lighting, hue, and atmosphere of the input image. In the base image feature prompt (804) presented as an example, the fifth sentence describes the features regarding the composition, and the sixth through eighth sentences describe the features regarding the person. The person described in the base image feature prompt (804) is an object that does not correspond to the incorrect answer area (802) and may correspond to an object appropriately identified by one or more artificial intelligence models mounted on the electronic device (1000).

[0191] According to one embodiment of the present disclosure, an electronic device (1000) may use a script writing tool to generate a base image feature prompt (804) written to fit a predetermined (e.g., specified, predetermined, (pre) determined, pre-set) template based on an input image. Meanwhile, according to one embodiment of the present disclosure, the electronic device (1000) may generate the base image feature prompt (804) through an artificial intelligence model (e.g., a generative model).

[0192] In step S821 of FIG. 8b, the server (2000) may generate a second feature prompt corresponding to an incorrect answer area (802) in the content image (801). In step S822 of FIG. 8b, the server (2000) may generate a training image (807) from the base image (805) based on the second feature prompt. In the present disclosure, the second feature prompt may also be referred to as an incorrect answer area image feature prompt (806).

[0193] In one embodiment of the present disclosure, the learning image (807) may correspond to an image materialized based on a second feature prompt corresponding to the incorrect answer area (802), from a base image (805) based on a first feature prompt corresponding to the remaining area excluding the incorrect answer area (802) among the content image (801).

[0194] In one embodiment of the present disclosure, the second feature prompt (or, incorrect answer area image feature prompt (806)) is created based on features extracted (or, described) from an input image, and the server (2000) can generate an image regarding an incorrect answer area (802) based on the second feature prompt (or, incorrect answer area image feature prompt (806)).

[0195] Referring to FIG. 8b and FIG. 8c together, a server (2000) according to one embodiment of the present disclosure can generate an incorrect answer area image feature prompt (806) through a prompt generation module (810). The server (2000) can generate an incorrect answer area image feature prompt (806) from an image extracted as an incorrect answer area (802) through a prompt generation module (810). By generating an incorrect answer area image feature prompt (806) from an image extracted as an incorrect answer area (802), the server (2000) can describe the incorrect answer area (802) in detail. For example, unlike generating a base image feature prompt (804), the prompt generation module (810) can extract (or describe) features regarding an object included in the incorrect answer area (802) in detail, thereby generating an incorrect answer area image feature prompt (806) in which the object is described in detail. For example, the prompt generation module (810) can generate an incorrect area image feature prompt (806) in which detailed features such as the type, size, color, composition, pattern, and pose of an object within an image are described from an input image.

[0196] The server (2000), through the prompt generation module (810), from the image of the incorrect answer area (802), [describes] "Type: Adolescent or young female," "Size: The overall height of the figure in the image occupies about 1 / 4 of the image, and her leg length and upper body proportions are naturally in harmony," "Color: She is wearing a gray hoodie and a dark red scarf, and her legs are exposed by a short checkered skirt. Her shoes are comfortable shoes such as dark-colored sneakers or loafers," "Composition: She is standing on the left side of the screen, turning her body slightly to the right to face a man, and is in a posture of handing a small green bouquet to the man in her right hand. Her gaze is naturally directed at the bouquet," "Pattern: There is no pattern on the top and scarf, and the skirt has a repeating dark-colored grid pattern (checkered pattern). Overall, the outfit is casual yet gives the impression of a school uniform," "Posture: She is extending her right hand forward to hand over the flower, and is standing in a stable posture with her legs slightly spread. The waist and An incorrect area image feature prompt (806) can be generated with the following content: "The upper body is naturally straightened and adopts a somewhat calm attitude," and "Other descriptions: The bob hair naturally falls to the sides of the ears, and the facial expression gives a peaceful and cautious impression. The scarf stands out as a contrasting element in the landscape, and the red color harmonizes with the surrounding calm colors." In the incorrect area image feature prompt (806) presented as the example above, features regarding the type, size, color, composition, pattern, posture, etc. of the object included in the incorrect area (802) are described in detail.

[0197] According to one embodiment of the present disclosure, by generating an incorrect answer area image feature prompt (806) separately from a base image feature prompt (804), the description of the incorrect answer area (802) can be made more detailed than when a feature prompt regarding the entire content image (801) is generated. Accordingly, the server (2000) can train an artificial intelligence model mounted on an electronic device (1000) using a training image (807) in which the incorrect answer area (802) is described in detail, and thereby, the identification performance of the incorrect answer area (802) that could not be identified previously can be improved.

[0198] Hereinafter, with reference to FIGS. 9a to 9c, the operation of the server (2000) that determines the correct answer data (or GT data) and identification difficulty of the generated training image will be explained in more detail.

[0199] FIG. 9a is a flowchart illustrating an exemplary operation for determining the identification difficulty of each of the content image and the learning image of a server (2000) according to one embodiment of the present disclosure. FIG. 9b is a diagram illustrating an exemplary operation for determining the recognition difficulty of each of the content image and the learning image of a server (2000) according to one embodiment of the present disclosure and storing the learning image. FIG. 9c is a diagram illustrating an exemplary operation for determining the recognition difficulty of each of the content image and the learning image of a server (2000) according to one embodiment of the present disclosure and storing the learning image.

[0200] Referring to FIGS. 9a through 9c, a server (2000) according to one embodiment of the present disclosure can determine the identification difficulty of each of the content image and the learning image. The server (2000) can determine the identification difficulty of the learning image based on the identification difficulty of the content image, and the learning image can be stored according to the identification difficulty.

[0201] In step S910 of FIG. 9a, the server (2000) can determine the identification difficulty of the content image as the first identification difficulty if the information corresponding to one or more objects identified from the content image through the first artificial intelligence model and the information corresponding to one or more objects identified from the content image through the artificial intelligence model stored in the server (2000) are identical. The server (2000) can determine the identification difficulty of the content image as the first identification difficulty based on the fact that the information corresponding to one or more objects identified from the content image through the first artificial intelligence model and the information corresponding to one or more objects identified from the content image through the artificial intelligence model stored in the server (2000) are identical.

[0202] Referring together with FIG. 9b, the electronic device (1000) may be equipped with a first artificial intelligence model (910) and a second artificial intelligence model (920), and the first artificial intelligence model (910) is a small model and the second artificial intelligence model (920) is a middle model, as illustrated in the example.

[0203] A first artificial intelligence model (910) mounted on an electronic device (1000) can identify one or more objects from an input content image (901) and output first object information (911) (or information corresponding to one or more objects identified through the first artificial intelligence model (910)) including an object class and an object location corresponding to the identified one or more objects. The first object information (911) includes information about one or more objects and can be displayed as "(object class, object location)".

[0204] For example, the first artificial intelligence model (910) can identify males and females from the input content image (901), and the first object information (911) regarding males can be indicated as (Boy, position of Boy), and the first object information (911) regarding females can be indicated as "(Girl, position of Girl)." For example, the object positions can be indicated by four numbers "(x, y, W, H)" representing a bounding box. x represents the x-coordinate of the top-left corner of the bounding box, y represents the y-coordinate of the top-left corner of the bounding box, W represents the width of the bounding box, and H represents the height of the bounding box. The position of Boy can be indicated as "(x1, y1, W1, H1)" and the position of Girl can be indicated as "(x2, y2, W2, H2)".

[0205] A second artificial intelligence model (920) mounted on an electronic device (1000) can identify one or more objects from an input content image (901) and output second object information (921) (or information corresponding to one or more objects identified through the second artificial intelligence model (920)) including an object class and an object location corresponding to the identified one or more objects. The second object information (921) includes information about one or more objects and can be displayed as "(object class, object location)".

[0206] For example, the second artificial intelligence model (920) can identify a male from the input content image (901), and the second object information (921) regarding the male can be displayed as "(Boy, location of Boy)". For example, the location of Boy can be displayed as "(x1, y1, W1, H1)".

[0207] The electronic device (1000) can identify that one or more objects identified through the first artificial intelligence model (910) and one or more objects identified through the second artificial intelligence model (920) are different from each other by comparing the first object information (911) obtained through the first artificial intelligence model (910) and the second object information (921) obtained through the second artificial intelligence model (920). For example, since a woman is identified through the first artificial intelligence model (910) but not through the second artificial intelligence model (920), the electronic device (1000) can identify that one or more objects identified through the first artificial intelligence model (910) and one or more objects identified through the second artificial intelligence model (920) are different from each other. Accordingly, the electronic device (1000) can transmit the content image (901) and the first object information (911) and second object information (921) regarding the content image to the server (2000).

[0208] In one embodiment of the present disclosure, a third artificial intelligence model (930) mounted on a server (2000) may identify one or more objects from a content image (901) received from an electronic device (1000) and output third object information (931) (or information corresponding to one or more objects identified through the third artificial intelligence model (930)) including an object class and an object location corresponding to the one or more identified objects. The third object information (931) includes information about one or more objects and may be indicated as "(object class, object location)".

[0209] For example, the third artificial intelligence model (930) can identify a male and a female from the content image (901), and the third object information (931) regarding the male can be indicated as "(Boy, location of Boy)" and the third object information (931) regarding the female can be indicated as "(Girl, location of Girl)". For example, the location of the Boy can be indicated as (x1, y1, W1, H1) and the location of the Girl can be indicated as "(x2, y2, W2, H2)".

[0210] The server (2000) can identify that one or more objects identified through the first artificial intelligence model (910) and one or more objects identified through the third artificial intelligence model (930) are identical by comparing the first object information (911) obtained through the first artificial intelligence model (910) and the third object information (931) obtained through the third artificial intelligence model (930). The server (2000) can identify that one or more objects identified through the second artificial intelligence model (920) and one or more objects identified through the third artificial intelligence model (930) are different by comparing the second object information (921) obtained through the second artificial intelligence model (920) and the third object information (931) obtained through the third artificial intelligence model (930). In this case, the electronic device (1000) can determine that the identification difficulty of the content image (901) has a first identification difficulty of a middle level.

[0211] The server (2000) may determine as an incorrect answer area one or more objects among those identified through the third artificial intelligence model (930) that are not identified through the second artificial intelligence model (920) or are identified as a different class from the second artificial intelligence model (920). For example, the server (2000) may determine a "Girl" that is not identified through the second artificial intelligence model (920) as an incorrect answer area.

[0212] Meanwhile, in step S920 of FIG. 9a, if the information corresponding to one or more objects identified from the content image through the second artificial intelligence model and the information corresponding to one or more objects identified from the content image through the artificial intelligence model stored in the server (2000) are identical, the server (2000) can determine the identification difficulty of the content image as the second identification difficulty. The server (2000) can determine the identification difficulty of the content image as the second identification difficulty based on the fact that the information corresponding to one or more objects identified from the content image through the second artificial intelligence model and the information corresponding to one or more objects identified from the content image through the artificial intelligence model stored in the server (2000) are identical.

[0213] Referring to FIG. 9a and FIG. 9c together, the first artificial intelligence model (910) identifies a male from the input content image (901), and the first object information (912) regarding the male is indicated as "(Boy, location of Boy)", for example, the location of the Boy can be indicated as "(x1, y1, W1, H1)". The second artificial intelligence model (920) identifies a male and a female from the input content image (901), and the second object information (922) regarding the male is indicated as "(Boy, location of Boy)" and the second object information (922) regarding the female is indicated as "(Girl, location of Girl)", for example, the location of the Boy can be indicated as "(x1, y1, W1, H1)" and the location of the Girl can be indicated as "(x2, y2, W2, H2)".

[0214] The electronic device (1000) can transmit the content image (901), the first object information (912), and the second object information (922) regarding the content image to the server (2000) as it identifies that one or more objects identified through the first artificial intelligence model (910) and one or more objects identified through the second artificial intelligence model (920) are different from each other.

[0215] The third artificial intelligence model (930) can identify males and females from content images (901) received from the electronic device (1000), and the third object information (932) regarding males can be displayed as "(Boy, location of Boy)" and the third object information (932) regarding females can be displayed as "(Girl, location of Girl)". For example, the location of Boy can be displayed as "(x1, y1, W1, H1)" and the location of Girl can be displayed as "(x2, y2, W2, H2)".

[0216] The server (2000) can identify that one or more objects identified through the second artificial intelligence model (920) and one or more objects identified through the third artificial intelligence model (930) are identical by comparing the second object information (922) obtained through the second artificial intelligence model (920) and the third object information (932) obtained through the third artificial intelligence model (930). The server (2000) can identify that one or more objects identified through the first artificial intelligence model (910) and one or more objects identified through the third artificial intelligence model (930) are different by comparing the first object information (912) obtained through the first artificial intelligence model (910) and the third object information (932) obtained through the third artificial intelligence model (930). In this case, the electronic device (1000) can determine that the identification difficulty of the content image (901) has a second identification difficulty of a hard level. The second identification difficulty can be higher than the first identification difficulty.

[0217] The server (2000) may determine as an incorrect answer area one or more objects identified through the third artificial intelligence model (930) that are not identified through the first artificial intelligence model (910) or are identified differently from the first artificial intelligence model (910). For example, the server (2000) may determine as an incorrect answer area an area corresponding to the bounding box of a "Girl" that is not identified through the first artificial intelligence model (910).

[0218] In step S930 of FIG. 9a, the server (2000) can determine the identification difficulty of the learning image as the first identification difficulty if the content image corresponding to the learning image has the first identification difficulty, and can determine the identification difficulty of the learning image as the second identification difficulty if the content image corresponding to the learning image has the second identification difficulty. The server (2000) can determine the identification difficulty of the learning image as the first identification difficulty based on the fact that the content image corresponding to the learning image has the first identification difficulty. The server (2000) can determine the identification difficulty of the learning image as the second identification difficulty based on the fact that the content image corresponding to the learning image has the second identification difficulty.

[0219] In one embodiment of the present disclosure, a content image (901) may include a first content image having a first identification difficulty based on the fact that information corresponding to one or more objects identified from the content image (901) through a first artificial intelligence model (910) and information corresponding to one or more objects identified from the content image (901) through a third artificial intelligence model (930) stored in a server (2000) are identical, and a second content image having a second identification difficulty based on the fact that information corresponding to one or more objects identified from the content image (901) through a second artificial intelligence model (920) and information corresponding to one or more objects identified from the content image (901) through a third artificial intelligence model (930) stored in a server (2000) are identical.

[0220] Referring to FIGS. 9b and FIGS. 9c together, the server (2000) can generate a training image (902) based on a content image (901) and an incorrect answer area extracted from the content image (901). Since the operation of generating the training image (902) has been explained with reference to FIGS. 8a through 8c, a detailed explanation will not be repeated here.

[0221] In one embodiment of the present disclosure, the server (2000) may identify one or more objects from a training image (902) through a third artificial intelligence model (930) and store information regarding one or more objects identified from the training image (902) as a response (or correct answer) to the object identification of the training image (902). That is, the server (2000) may identify one or more objects from the training image (902) through the third artificial intelligence model (930) in order to obtain ground-truth (GT) data (904, 906) regarding the training image (902).

[0222] In the present disclosure, ground-truth (GT) data may be expressed in various ways, such as actual data of a data sample, actual data, observed data, ground truth data, ground truth information, observed information, observed information, actual data, or labeled results. GT data may be information included in the data sample, information corresponding to the data sample, information associated with the data sample, or information mapped to the data sample. GT data may be associated with, corresponding to, or mapped to one or more labels.

[0223] In one embodiment of the present disclosure, GT data (904, 906) may be represented as (object class, object location, object confidence). A server (2000) may obtain class, location, and confidence information of one or more objects identified from a training image (902) through a third artificial intelligence model (930). The confidence information represents the probability that an object belongs to a corresponding class and may be expressed as a value between 0 and 1.

[0224] For example, the third artificial intelligence model (930) can identify males and females from the input training image (902), and object information regarding males can be displayed as "(Boy, location of Boy, confidence level of Boy)", and object information regarding females can be displayed as "(Girl, location of Girl, confidence level of Girl)". For example, the location of Boy can be displayed as "(x1', y1', W1', H1')", and the location of Girl can be displayed as "(x2', y2', W2', H2')". For example, the confidence level of Boy, C1, can be displayed as a value from 0 to 1, and the confidence level of Girl, C2, can be displayed as a value from 0 to 1.

[0225] As illustrated in FIG. 9b, if the content image (901) corresponding to the training image (902) has a first identification difficulty (e.g., Middle), the identification difficulty of the training image (902) can also be determined to have a first identification difficulty (e.g., Middle). In this case, the server (2000) can store the training image (902) determined to have a first identification difficulty in a first identification difficulty dataset (903) (or a Middle dataset).

[0226] On the other hand, as illustrated in FIG. 9c, if the content image (901) corresponding to the training image (902) has a second identification difficulty (e.g., Hard), the identification difficulty of the training image (902) can also be determined to have a second identification difficulty (e.g., Hard). In this case, the server (2000) can store the training image (902) determined to have a second identification difficulty in a second identification difficulty dataset (905) (or a Hard dataset).

[0227] In one embodiment of the present disclosure, the learning image (902) may include a first learning image corresponding to a first content image having a first identification difficulty and a second learning image corresponding to a second content image having a second identification difficulty.

[0228] Hereinafter, with reference to FIGS. 10a to 10c, the operation of a server (2000) that learns artificial intelligence models mounted on an electronic device (1000) using a stored dataset will be explained in more detail below.

[0229] FIG. 10a is a flowchart illustrating an exemplary operation of learning a model of a server (2000) according to one embodiment of the present disclosure. FIG. 10b is a flowchart illustrating an exemplary operation of learning a model of a server (2000) according to one embodiment of the present disclosure. FIG. 10c is a diagram illustrating an exemplary operation of learning a model of a server (2000) according to one embodiment of the present disclosure.

[0230] Step S1010 of FIG. 10a is a step that embodies Step S340 of FIG. 3. Step S1020 of FIG. 10b is a step that embodies Step S340 of FIG. 3.

[0231] In step S1010 of FIG. 10a, the server (2000) can learn a model corresponding to a first artificial intelligence model based on each of a first type dataset and a second type dataset, in which the ratio between a training image having a first identification difficulty and a training image having a second identification difficulty is different.

[0232] In one embodiment of the present disclosure, an electronic device (1000) may receive from a server (2000) either a first type artificial intelligence model and a second type artificial intelligence model trained based on each of a first type dataset and a second type dataset, wherein the ratio between a first training image and a second training image is different from each other based on a first artificial intelligence model.

[0233] In step S1020 of FIG. 10b, the server (2000) can learn a model corresponding to the second artificial intelligence model based on a third type dataset and a fourth type dataset, respectively, in which the ratio between a training image having a first identification difficulty and a training image having a second identification difficulty is different. That is, the server (2000) can learn the first artificial intelligence model and the second artificial intelligence model, respectively, using training images classified by the first identification difficulty or the second identification difficulty.

[0234] In one embodiment of the present disclosure, an electronic device (1000) may receive from a server (2000) either a third type artificial intelligence model or a fourth type artificial intelligence model trained based on each of a third type dataset and a fourth type dataset, wherein the ratio between a first training image and a second training image is different from each other.

[0235] Referring together to FIGS. 10a to 10c, in one embodiment of the present disclosure, a server (2000) can determine the identification difficulty of the generated training image after generating a training image. Based on the identification difficulty, the server (2000) can classify and store the training image into a first identification difficulty dataset (1001) (e.g., Hard dataset) or a second identification difficulty dataset (1002) (e.g., Middle dataset).

[0236] In one embodiment of the present disclosure, the server (2000) may compare (1003) the amount of a newly collected dataset and the amount of an existing training dataset before performing model training. In the present disclosure, 'existing training dataset' refers to a training dataset used for prior training before the first and second artificial intelligence models were first deployed to the electronic device (1000). In the present disclosure, 'newly collected dataset' may refer to a training dataset stored in the server (2000) for updating the first and second artificial intelligence models after the first and second artificial intelligence models have been deployed to the electronic device (1000). In one embodiment of the present disclosure, the 'newly collected dataset' may include training images and GT data newly generated in the server (2000).

[0237] The server (2000) may not perform model training when the amount of the new collected dataset is less than the amount of the existing training dataset, and may acquire the new collected dataset until the amount of the new collected dataset becomes greater than the amount of the existing training dataset. The server (2000) may perform model training using the new collected dataset when the amount of the new collected dataset is greater than or equal to the amount of the existing training dataset. Considering that when the amount of the new collected dataset is less than the amount of the existing training dataset, the degree of improvement in model performance may be minimal even if the model is trained using a small amount of the new collected dataset, the model may only be trained when the amount of the new collected dataset is greater than or at least equal to the amount of the existing training dataset.

[0238] In one embodiment of the present disclosure, a server (2000) can train a first artificial intelligence model using a newly collected dataset. To train the first artificial intelligence model, the server (2000) can set multiple types of datasets by varying the ratio of a first identification difficulty dataset (1001) (e.g., a hard dataset) and a second identification difficulty dataset (1002) (e.g., a middle dataset) stored in the newly collected dataset.

[0239] For example, the server (2000) may configure the first-1 type dataset (1011) as '40% Hard dataset and 60% Middle dataset', the second-1 type dataset (1012) as '60% Hard dataset and 40% Middle dataset', and the N-1 type dataset (1013) (N is a natural number greater than or equal to 3) as '90% Hard dataset and 10% Middle dataset'. FIG. 10c illustrates that there are three or more types of datasets for training the first artificial intelligence model, but this is exemplary and the number of types of datasets is not limited thereto. Meanwhile, in the present disclosure, two type datasets selected arbitrarily among the 1-1 type dataset (1011) to the N-1 type dataset (1013) may be referred to as the 1st type dataset and the 2nd type dataset, respectively.

[0240] In one embodiment of the present disclosure, the server (2000) can train a second artificial intelligence model using a newly collected dataset. To train the second artificial intelligence model similarly to when training the first artificial intelligence model, the server (2000) can set multiple types of datasets by varying the ratio of the first identification difficulty dataset (1001) (e.g., Hard dataset) and the second identification difficulty dataset (1002) (e.g., Middle dataset) stored in the newly collected dataset.

[0241] For example, the server (2000) may configure the first-2 type dataset (1021) as '40% Hard dataset and 60% Middle dataset', the second-2 type dataset (1022) as '60% Hard dataset and 40% Middle dataset', and the N-2 type dataset (1023) (N is a natural number greater than or equal to 3) as '90% Hard dataset and 10% Middle dataset'. FIG. 10c illustrates that there are three or more types of datasets for training the second artificial intelligence model, but this is exemplary and the number of types of datasets is not limited thereto. Meanwhile, in the present disclosure, any two type datasets among the first-2 type dataset (1021) to the N-2 type dataset (1023) may be referred to as the third type dataset and the fourth type dataset, respectively.

[0242] In one embodiment of the present disclosure, if it is determined that an update of the first artificial intelligence model and / or the second artificial intelligence model is necessary as a result of model learning, the server (2000) may transmit information corresponding to the update of the first artificial intelligence model and / or the second artificial intelligence model to the electronic device (1000).

[0243] For example, when training a first artificial intelligence model, the server (2000) may perform a first-1 type model training (1031) based on a first-1 type dataset (1011), perform a second-1 type model training (1032) based on a second-1 type dataset (1012), and perform an N-1 type model training (1033) based on an N-1 type dataset (1013).

[0244] In one embodiment of the present disclosure, a server (2000) may obtain the first artificial intelligence model itself updated as a result of the first-1 type model learning (1031), the first artificial intelligence model itself updated as a result of the second-1 type model learning (1032), and the first artificial intelligence model itself updated as a result of the N-1 type model learning (1033). In the present disclosure, any two models among the first artificial intelligence model itself updated as a result of the first-1 type model learning (1031), the first artificial intelligence model itself updated as a result of the second-1 type model learning (1032), and the first artificial intelligence model itself updated as a result of the N-1 type model learning (1033) may be referred to as the first type artificial intelligence model and the second type artificial intelligence model, respectively.

[0245] In one embodiment of the present disclosure, the server (2000) may store the first artificial intelligence model itself updated as a result of the first-1 type model learning (1031), the first artificial intelligence model itself updated as a result of the second-1 type model learning (1032), and the first artificial intelligence model itself updated as a result of the N-1 type model learning (1033) in the model storage (1004). The server (2000) may distribute the model having the highest performance among the first artificial intelligence models updated in each type to the electronic device (1000).

[0246] In one embodiment of the present disclosure, a server (2000) may acquire parameter information updated as a result of a first-1 type model learning (1031), parameter information updated as a result of a second-1 type model learning (1032), and parameter information updated as a result of an N-1 type model learning (1033), and store them in a model storage (1004). The server (2000) may transmit information having high performance among the parameter information updated in each type to an electronic device (1000). The electronic device (1000) may receive the updated parameter information and update the first artificial intelligence model.

[0247] For example, when training a second artificial intelligence model, the server (2000) may perform a first-2 type model training (1041) based on a first-2 type dataset (1021), perform a second-2 type model training (1042) based on a second-2 type dataset (1022), and perform an N-1 type model training (1043) based on an N-1 type dataset (1023).

[0248] In one embodiment of the present disclosure, a server (2000) may obtain the second artificial intelligence model itself updated as a result of the first-2 type model learning (1041), the second artificial intelligence model itself updated as a result of the second-2 type model learning (1042), and the second artificial intelligence model itself updated as a result of the N-2 type model learning (1043). In the present disclosure, any two models among the second artificial intelligence model itself updated as a result of the first-2 type model learning (1041), the second artificial intelligence model itself updated as a result of the second-2 type model learning (1042), and the second artificial intelligence model itself updated as a result of the N-2 type model learning (1043) may be referred to as the third type artificial intelligence model and the fourth type artificial intelligence model, respectively.

[0249] The method by which the server (2000) learns the second artificial intelligence model and transmits information regarding the update of the second artificial intelligence model is similar to the method by which the server (2000) learns the first artificial intelligence model and transmits information regarding the update of the first artificial intelligence model, so a detailed explanation thereof is not duplicated here.

[0250] Hereinafter, with reference to FIGS. 11a and FIGS. 11b, a method for updating an artificial intelligence model mounted on an electronic device (1000) after model training will be explained in more detail below.

[0251] FIG. 11a is a flowchart illustrating an exemplary operation of distributing a model to a plurality of electronic devices (1000) of a server (2000) according to one embodiment of the present disclosure. FIG. 11b is a diagram illustrating an exemplary operation of distributing a model to a plurality of electronic devices (1000) of a server (2000) according to one embodiment of the present disclosure.

[0252] In step S1110 of FIG. 11a, the server (2000) may store a first type artificial intelligence model trained based on a first type dataset and a second type artificial intelligence model trained based on a second type dataset. Since the description of step S1110 has been explained with reference to FIG. 10b, a detailed description will not be repeated here.

[0253] In step S1120 of FIG. 11a, the server (2000) may deploy a first type artificial intelligence model to some of at least one electronic device (1000) and deploy a second type artificial intelligence model to other of at least one electronic device (1000).

[0254] In one embodiment of the present disclosure, at least one electronic device (1000) may receive from a server (2000) either a first type artificial intelligence model and a second type artificial intelligence model trained based on each of a first type dataset and a second type dataset having different ratios between a first training image and a second training image based on a first artificial intelligence model.

[0255] Referring together with FIG. 11b, the server (2000) can perform data communication with a plurality of electronic devices (1000). Before transmitting information regarding an update (e.g., distribution of an updated model) to all electronic devices (1000) communicating with the server (2000), the server (2000) can perform tests of candidate models through some of the electronic devices. The server (2000) can distribute candidate models stored in the model storage (1111) to some of the electronic devices.

[0256] For example, the server (2000) may distribute a first type artificial intelligence model (1101) trained based on a first type dataset to 10% of the electronic devices (1000a) (hereinafter referred to as the first group of electronic devices (1000a)). The first group of electronic devices (1000a) may test (1103) object recognition performance for input content images using the distributed first type artificial intelligence model (1101). The server (2000) may distribute a second type artificial intelligence model (1102) trained based on a second type dataset to another 10% of the electronic devices (1000b) (hereinafter referred to as the second group of electronic devices (1000b)). The second group of electronic devices (1000b) can test the object recognition performance (1104) for input content images using the distributed second type artificial intelligence model (1102). The remaining electronic devices (1000c), excluding the first group and second group of electronic devices (1000a, 1000b) among all electronic devices (1000), can test the object recognition performance (1105) for input content images using the previously installed first artificial intelligence model.

[0257] In one embodiment of the present disclosure, at least one electronic device (1000) can transmit (or send) the test result of either a first type artificial intelligence model (1101) and a second type artificial intelligence model (1102) to a server (2000).

[0258] In step S1130 of FIG. 11a, the server (2000) may receive information regarding the test results of a first type artificial intelligence model from some of at least one electronic device (1000). In step S1140 of FIG. 11a, the server (2000) may receive information regarding the test results of a second type artificial intelligence model from another of at least one electronic device (1000). In step S1150 of FIG. 11a, the server (2000) may receive information regarding the test results of the first artificial intelligence model from the remainder of at least one electronic device (1000), excluding some and other parts.

[0259] Referring to FIG. 11a and FIG. 11b together, the first group of electronic devices (1000a) can transmit test result data (1106) of the first type of artificial intelligence model (1101) to the server (2000), the second group of electronic devices (1000b) can transmit test result data (1107) of the second type of artificial intelligence model (1102) to the server (2000), and the remaining electronic devices (1000c) can transmit test result data (1108) of the existing model to the server (2000). For example, the server (2000) may receive test result data (1106) of a first type artificial intelligence model (1101) from a first group of electronic devices (1000a), receive test result data (1107) of a second type artificial intelligence model (1102) from a second group of electronic devices (1000b), and receive test result data (1108) of an existing model from the remaining electronic devices (1000c). For example, the test result data may include an Intersection over Union (IoU), which is an indicator measuring how much the predicted bounding box and the actual bounding box overlap; Precision, which represents the ratio of the actual object among the objects predicted by the model; Recall, which represents the ratio of the object recognized by the model among the objects that actually exist; and mean Average Precision (mAP), which is an indicator that comprehensively evaluates the precision and recall of the model.

[0260] In step S1160 of FIG. 11a, the server (2000) can determine the final model among the first artificial intelligence model, the first artificial intelligence model, and the second artificial intelligence model based on information regarding the test results of the first type artificial intelligence model, information regarding the test results of the second type artificial intelligence model, and information regarding the test results of the first artificial intelligence model.

[0261] In step S1170 of FIG. 11a, if the determined final model is a first type artificial intelligence model or a second type artificial intelligence model, the server (2000) can distribute the determined first type artificial intelligence model or the second type artificial intelligence model to at least one electronic device (1000). Based on the fact that the determined final model is a first type artificial intelligence model or a second type artificial intelligence model, the server (2000) can distribute the determined first type artificial intelligence model or the second type artificial intelligence model to at least one electronic device (1000).

[0262] In one embodiment of the present disclosure, at least one electronic device (1000) may receive from a server (2000) information corresponding to a final model determined based on the test results of either the first type artificial intelligence model and the second type artificial intelligence model among the first artificial intelligence model, the first type artificial intelligence model, and the second type artificial intelligence model. For example, at least one electronic device (1000) may receive from a server (2000) information corresponding to a final model determined based on the test results of the first artificial intelligence model, the test results of the first type artificial intelligence model, and the test results of the second type artificial intelligence model among the first artificial intelligence model, the first type artificial intelligence model, and the second type artificial intelligence model. In one embodiment of the present disclosure, at least one electronic device (1000) may update the first artificial intelligence model based on the information corresponding to the received final model.

[0263] Referring to FIG. 11b together, for example, the server (2000) can determine that the test result of the first type artificial intelligence model (1101) has the best performance by referring to the first type model test result data (1106), the second type model test result data (1107), and the existing model test result data (1108). Accordingly, the server (2000) can determine (1109) the first type artificial intelligence model (1101) as the final model.

[0264] In one embodiment of the present disclosure, the server (2000) can determine (1110) whether the determined final model is a first type artificial intelligence model (1101) or a second type artificial intelligence model (1102). If the determined final model is a first type artificial intelligence model (1101) or a second type artificial intelligence model (1102), the server (2000) can distribute the final model to all electronic devices (1000). Meanwhile, if the determined final model is an existing model that is not a first type artificial intelligence model (1101) or a second type artificial intelligence model (1102), the server (2000) can delete the first type artificial intelligence model (1101) and the second type artificial intelligence model (1102) newly stored in the model storage (1111).

[0265] FIG. 11b illustrates, by way of example, that the server (2000) distributes the updated model itself, but the present disclosure is not limited thereto, and the server (2000) may transmit updated parameter information, and the electronic device (1000) may update the parameters of the previously installed model based on the received parameter information.

[0266] FIGS. 11a and 11b describe a method for updating an artificial intelligence model mounted on an electronic device (1000) based on a first artificial intelligence model, but a similar method can be applied to updating a second artificial intelligence model.

[0267] In one embodiment of the present disclosure, the server (2000) may store a third type artificial intelligence model learned based on a third type dataset and a fourth type artificial intelligence model learned based on a fourth type dataset.

[0268] In one embodiment of the present disclosure, the server (2000) may deploy a third type artificial intelligence model to some of at least one electronic device (1000) and deploy a fourth type artificial intelligence model to other of at least one electronic device (1000).

[0269] In one embodiment of the present disclosure, at least one electronic device (1000) may receive from a server (2000) either a third type artificial intelligence model and a fourth type artificial intelligence model trained based on each of a third type dataset and a fourth type dataset, wherein the ratio between a third training image and a fourth training image is different based on a second artificial intelligence model.

[0270] In one embodiment of the present disclosure, a server (2000) may receive information regarding the test results of a third type artificial intelligence model from some of at least one electronic device (1000). In one embodiment of the present disclosure, a server (2000) may receive information regarding the test results of a fourth type artificial intelligence model from another of at least one electronic device (1000). In one embodiment of the present disclosure, a server (2000) may receive information regarding the test results of a second artificial intelligence model from the remainder of at least one electronic device (1000), excluding some of the and other parts.

[0271] In one embodiment of the present disclosure, the server (2000) can determine the final model among the second artificial intelligence model, the third artificial intelligence model, and the fourth artificial intelligence model based on information regarding the test results of the third type artificial intelligence model, information regarding the test results of the fourth type artificial intelligence model, and information regarding the test results of the second artificial intelligence model.

[0272] In one embodiment of the present disclosure, if the determined final model is a third type artificial intelligence model or a fourth type artificial intelligence model, the server (2000) can distribute the determined third type artificial intelligence model or the fourth type artificial intelligence model to at least one electronic device (1000).

[0273] In one embodiment of the present disclosure, at least one electronic device (1000) may receive from a server (2000) information corresponding to a final model determined based on the test results of either the third type artificial intelligence model or the fourth type artificial intelligence model among the second artificial intelligence model, the third type artificial intelligence model, and the fourth type artificial intelligence model. For example, at least one electronic device (1000) may receive from a server (2000) information corresponding to a final model determined based on the test results of the second artificial intelligence model, the test results of the third type artificial intelligence model, and the test results of the fourth type artificial intelligence model among the second artificial intelligence model, the third type artificial intelligence model, and the fourth type artificial intelligence model. In one embodiment of the present disclosure, at least one electronic device (1000) may update the second artificial intelligence model based on the information corresponding to the received final model.

[0274] FIG. 12a is a flowchart illustrating an exemplary operation of learning a model of an electronic device (1000) according to one embodiment of the present disclosure. FIG. 12b is a diagram illustrating an exemplary operation of learning a model of an electronic device (1000) according to one embodiment of the present disclosure.

[0275] In step S1210 of FIG. 12a, an electronic device (1000) according to one embodiment of the present disclosure may store information corresponding to a content image and one or more objects identified through a second artificial intelligence model if one or more objects identified through a first artificial intelligence model and one or more objects identified through a second artificial intelligence model do not correspond to each other. The electronic device (1000) may store information corresponding to a content image and one or more objects identified through a second artificial intelligence model based on the fact that one or more objects identified through a first artificial intelligence model and one or more objects identified through a second artificial intelligence model do not correspond to each other.

[0276] Referring together with FIG. 12b, in one embodiment of the present disclosure, an object identification module (1210) within an electronic device (1000) may include a first artificial intelligence model corresponding to a small model (1211) and a second artificial intelligence model corresponding to a middle model (1212). Hereinafter, the description will be based on the assumption that the first artificial intelligence model is the small model (1211) and the second artificial intelligence model is the middle model (1212). At the time of operation in which the electronic device (1000) provides content images, it may perform object identification from a plurality of input content images using the small model (1211) and perform object identification from a plurality of input content images using the middle model (1212).

[0277] If the object identification result in the small model (1211) and the object identification result in the middle model (1212) differ with respect to specific content images, the electronic device (1000) may store some of the specific content images (e.g., 80%) in local storage (1220) within the electronic device (1000) and transmit the remainder of the specific content images (e.g., 20%) to a server (2000).

[0278] The electronic device (1000) can store some of the specific content images and object information identified from the some images through the middle model (1212) in the local storage (1220) within the electronic device (1000). The electronic device (1000) can transmit the remainder of the specific content images, object information identified from the remainder images through the small model (1211), and object information identified from the remainder images through the middle model (1212) to the server (2000). The object information may include object class information and object location information.

[0279] In step S1220 of FIG. 12a, the electronic device (1000) can learn the first artificial intelligence model by inputting information corresponding to one or more objects identified through the second artificial intelligence model as a response (or correct answer) to the object identification of the content image.

[0280] Referring together with FIG. 12b, the electronic device (1000) can learn a small model (1211) while not performing operations such as providing images. The electronic device (1000) can learn a small model (1211) using content images stored in local storage (1220). The electronic device (1000) can learn a small model (1211) by inputting object information identified through a middle model (1212) stored in local storage (1220) as the response (or correct answer) of the corresponding content image.

[0281] If the electronic device (1000) determines that an update to the small model (1211) is necessary during the process of learning the small model (1211), it may transmit information regarding the update to the small model (1211). For example, the electronic device (1000) may transmit updated parameter information to the small model (1211) or transmit the updated model itself.

[0282] According to one embodiment of the present disclosure, as the small model (1211) is learned within the electronic device (1000), a large amount of data may not be transmitted to the server (2000) for model learning. Additionally, as the small model (1211) is learned within the electronic device (1000), the electronic device (1000) may learn the small model (1211) by using the input content images as they are.

[0283] FIG. 13 is a block diagram showing an exemplary configuration of an electronic device (1000) according to one embodiment of the present disclosure.

[0284] Referring to FIG. 13, an electronic device (1000) according to one embodiment of the present disclosure may include a communication interface (110) (e.g., including a communication circuit), a processor (120) (e.g., including a processing circuit), a memory (130), a display (140), a video processing unit (145) (e.g., including various circuits and / or executable program instructions), a tuner unit (160), an audio processing unit (155) (e.g., including various circuits and / or executable program instructions), an audio output unit (150) (e.g., including a circuit), a sensing unit (170) (e.g., including a circuit), an input / output unit (180) (e.g., including various circuits), and a user input unit (190) (e.g., including a user interface circuit). However, not all components illustrated in FIG. 13 are essential components. The electronic device (1000) may be implemented with more components than those shown in FIG. 13, or with fewer components.

[0285] The memory (130) can store instructions, algorithms, data structures, program codes, and application programs that are stored for processing and control of the processor (120), and can store data that is input to or output from the electronic device (1000). The memory (130) may include at least one of, for example, a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), RAM (Random Access Memory), SRAM (Static Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), PROM (Programmable Read-Only Memory), Mask ROM, Flash ROM, etc.), a hard disk drive (HDD), or a solid-state drive (SSD). A program (one or more instructions) or application stored in memory (130) can be executed by the processor (120).

[0286] The tuner unit (160) can select only the frequency of the channel to be received by the electronic device (1000) from among many radio wave components by tuning through amplification, mixing, resonance, etc. of broadcast content received via wired or wireless connection. The broadcast signal received through the tuner unit (160) is separated into audio, video, and additional information (e.g., EPG (Electronic Program Guide)). The separated audio, video, and additional information can be stored in memory (130) under the control of the processor (120).

[0287] The tuner unit (160) can receive broadcast signals from various sources such as terrestrial broadcasting, cable broadcasting, satellite broadcasting, internet broadcasting, etc. The tuner unit (160) can also receive broadcast signals from sources such as analog broadcasting or digital broadcasting.

[0288] The communication interface (110) may include various communication circuits and, under the control of the processor (120), may connect the electronic device (1000) to peripheral devices, external devices, servers, display devices, remote control devices, mobile terminals, etc. The communication interface (110) may include at least one communication module capable of performing wireless communication. For example, the communication interface (110) may separately provide a communication module that communicates with a server, a communication module that communicates with a display device, a communication module that communicates with a remote control device, and a communication module that communicates with a mobile terminal, or it may include a single integrated module.

[0289] The communication interface (110) may include at least one of a wireless LAN module (111), a Bluetooth module (112), and a wired Ethernet (113) depending on the performance and structure of the electronic device (1000). The Bluetooth module (112) can receive Bluetooth signals transmitted from a peripheral device according to the Bluetooth communication standard. The Bluetooth module (112) can be a BLE (Bluetooth Low Energy) communication module and can receive BLE signals. The Bluetooth module (112) can scan for BLE signals continuously or temporarily to detect whether a BLE signal is being received. The wireless LAN module (111) can transmit and receive Wi-Fi signals with a peripheral device according to the Wi-Fi communication standard.

[0290] The sensing unit (or sensing interface, 170) may include various circuits and / or executable program instructions, detect the user's voice, the user's image, or the user's interaction, and may include a microphone (171), a sensor (172), and an optical receiver (173).

[0291] The microphone (171) can receive an audio signal including the user's uttered voice or noise and can convert the received audio signal into an electrical signal and output it to the processor (120).

[0292] The microphone (171) may be provided in a remote control device such as a remote control, a mobile terminal, or an AI speaker. For example, the mobile terminal may run an application to remotely control the electronic device (1000). In this case, the microphone (171) provided in the remote control device may receive an audio signal containing the user's uttered voice or noise. The remote control device may convert the audio signal into a control signal and transmit it to the electronic device (1000). The electronic device (1000) may receive the control signal from the remote control device through the communication unit (110).

[0293] In one embodiment of the present disclosure, an electronic device (1000) may transmit a received voice signal to an external server (e.g., a speech-to-text (STT) server). The external server may generate text from the content of the received voice signal. The external server may transmit text information corresponding to the user's voice back to the electronic device or to another server. The electronic device may receive text information corresponding to the user's voice from the external server. Meanwhile, the present disclosure is not limited thereto, and the electronic device (1000) may convert a received voice signal into text within the electronic device (1000) to generate text information corresponding to the user's voice. The electronic device (1000) may directly use the text information generated internally or transmit it to another external server.

[0294] The sensor (172) detects the user's image, or the user's interaction, gesture, and touch, and may include a distance sensor, an image sensor, a gesture sensor, an illuminance sensor, etc. The distance sensor may include various sensors that detect the distance between the electronic device (1000) and the user, such as an ultrasonic sensor, an IR (Infrared Radiation) sensor, or a TOF (Time Of Flight) sensor. The distance sensor detects the distance to the user and can transmit the sensing data to the processor (120). The image sensor can capture the user's gesture by photographing it through a camera, etc., and transmit it to the processor (120). The gesture sensor can detect the speed or direction of movement through an accelerometer or a gyroscope. The illuminance sensor can detect ambient illuminance.

[0295] The optical receiver (173) may include various circuits and may receive an optical signal (including a control signal). The optical receiver (173) may receive an optical signal corresponding to user input (e.g., touch, press, touch gesture, voice, or motion) from a control device such as a remote control or a mobile phone.

[0296] The input / output unit (180) can receive video (e.g., dynamic image signal or still image signal), audio (e.g., voice signal or music signal), and additional information from an external device, etc., under the control of the processor (120). The input / output unit (180) may include a port for outputting video and audio together, and may also include separate ports for outputting video and audio separately.

[0297] The input / output unit (180) may include various circuits and may include one of an HDMI port (High-Definition Multimedia Interface port, 181), a component jack (182), a PC port (183), and a USB port (184). The input / output unit (180) may include a combination of an HDMI port (181), a component jack (182), a PC port (183), and a USB port (184). Additionally, the input / output unit (180) may include one of a DP (Display Port), a Thunderbolt port, a VGA (Video Graphics Array) port, an RGB port, a D-SUB, and a DVI (Digital Visual Interface).

[0298] When the electronic device (1000) corresponds to a content providing device such as a set-top box, the input / output unit (180) can output video, audio, and additional information to a display device under the control of the processor (120).

[0299] In one embodiment of the present disclosure, image data and voice data are transmitted through separate ports within the input / output unit (180) and may be stored as separate tracks in the electronic device (1000). For example, image data may be transmitted through ports such as VGA, DVI, etc., and voice data may be transmitted through separate ports. Alternatively, image data and voice data may be transmitted as a single stream through HDMI, DP, Thunderbolt, etc., and may be stored as separate tracks in the electronic device (1000).

[0300] The video processing unit (145) may include various circuits and / or executable program instructions, process image data to be displayed by the display (140), and perform various image processing operations such as decoding, rendering, scaling, noise filtering, frame rate conversion, and resolution conversion on the image data.

[0301] The display (140) can output content received from a broadcasting station or from an external device such as an external server or an external storage medium. The content is a media signal and may include a video signal, an image, a text signal, etc.

[0302] The audio processing unit (155) may include various circuits and / or executable program instructions and may perform processing on audio data. Various processing such as decoding, amplification, and noise filtering on audio data may be performed in the audio processing unit (155).

[0303] The audio output unit (150) may include various circuits and may output audio included in content received through the tuner unit (160) under the control of the processor (120), audio input through the communication unit (110) or input / output unit (180), and audio stored in memory (130). The audio output unit (150) may include at least one of a speaker (151), headphones (152), or S / PDIF (Sony / Philips Digital Interface: output terminal) (153).

[0304] The user input unit (190) may include various user interface circuits and may receive user input for controlling the electronic device (1000). The user input unit (190) may include, but is not limited to, various forms of user input devices including a touch panel that detects a user's touch, a button that receives a user's push operation, a wheel that receives a user's rotation operation, a keyboard, a dome switch, a microphone for voice recognition, a motion detection sensor that senses motion, etc. When a remote control device, such as a remote control device or other mobile terminal, controls the electronic device (1000), the user input unit (190) may receive a control signal received from the remote control device.

[0305] According to one embodiment of the present disclosure, an electronic device (1000) is provided.

[0306] According to one embodiment of the present disclosure, an electronic device (1000) may include a communication interface (110), a memory (130) for storing a plurality of instructions, and at least one processor (120) operably coupled to the memory (130) and comprising a processing circuitry.

[0307] According to one embodiment of the present disclosure, by having at least one processor (120) execute instructions alone or in cooperation, the electronic device (1000) can identify one or more objects through each of a first artificial intelligence model and a second artificial intelligence model based on a content image input to the electronic device (1000). According to one embodiment of the present disclosure, by having at least one processor (120) execute instructions alone or in cooperation, the electronic device (1000) can transmit to a server, through a communication interface (110), a content image, information corresponding to one or more objects identified through the first artificial intelligence model, and information corresponding to one or more objects identified through the second artificial intelligence model, based on the fact that one or more objects identified through the first artificial intelligence model from the content image and one or more objects identified through the second artificial intelligence model from the content image do not correspond to each other. According to one embodiment of the present disclosure, by having at least one processor (120) execute instructions alone or in cooperation, the electronic device (1000) can receive from a server information corresponding to an update of at least one of the first artificial intelligence model or the second artificial intelligence model, which is obtained using a training image generated based on a content image, information corresponding to one or more objects identified through a first artificial intelligence model, and information corresponding to one or more objects identified through a second artificial intelligence model, through a communication interface (110).

[0308] According to one embodiment of the present disclosure, by having at least one processor (120) execute instructions alone or in cooperation, the electronic device (1000) can receive an input corresponding to object identification for each of a plurality of content images. According to one embodiment of the present disclosure, by having at least one processor (120) execute instructions alone or in cooperation, the electronic device (1000) can execute a first artificial intelligence model and a second artificial intelligence model respectively upon receiving an input corresponding to object identification based on the fact that the resolution of a plurality of content images is less than a threshold value, and execute the first artificial intelligence model upon receiving an input corresponding to object identification based on the fact that the resolution of a plurality of content images is greater than or equal to a threshold value, and execute the second artificial intelligence model in response to a predetermined frequency.

[0309] According to one embodiment of the present disclosure, by having at least one processor (120) execute instructions alone or in cooperation, the electronic device (1000) can receive an input corresponding to object identification for each of a plurality of content images. According to one embodiment of the present disclosure, by having at least one processor (120) execute instructions alone or in cooperation, the electronic device (1000) can alternately execute a first artificial intelligence model and a second artificial intelligence model upon receiving an input corresponding to object identification. According to one embodiment of the present disclosure, by having at least one processor (120) execute instructions alone or in cooperation, the electronic device (1000) can execute both the first artificial intelligence model and the second artificial intelligence model when the frequency of execution of the first artificial intelligence model and the second artificial intelligence model corresponds to a predetermined frequency.

[0310] According to one embodiment of the present disclosure, by having at least one processor (120) execute instructions alone or in cooperation, the electronic device (1000) can store information corresponding to a content image and one or more objects identified through a second artificial intelligence model, provided that one or more objects identified through a first artificial intelligence model and one or more objects identified through a second artificial intelligence model do not correspond to each other. According to one embodiment of the present disclosure, by having at least one processor (120) execute instructions alone or in cooperation, the electronic device (1000) can learn the first artificial intelligence model by inputting information corresponding to one or more objects identified through a second artificial intelligence model as a response to object identification of the content image.

[0311] According to one embodiment of the present disclosure, information corresponding to an update of at least one of a first artificial intelligence model and a second artificial intelligence model may include at least one of a model updated based on information corresponding to a training image and one or more objects identified from the training image via a server based on the first artificial intelligence model, or a model updated based on information corresponding to a training image and one or more objects identified from the training image via a server based on the second artificial intelligence model.

[0312] According to one embodiment of the present disclosure, the learning image may be based on an incorrect answer area and a content image corresponding to an object differently identified between the electronic device (1000) and the server.

[0313] According to one embodiment of the present disclosure, a learning image may correspond to an image materialized based on a second feature prompt corresponding to an incorrect answer area, from a base image based on a first feature prompt corresponding to the remaining area excluding the incorrect answer area among the content images.

[0314] According to one embodiment of the present disclosure, a content image may include a first content image having a first identification difficulty based on the fact that information corresponding to one or more objects identified from the content image through a first artificial intelligence model and information corresponding to one or more objects identified from the content image through an artificial intelligence model stored in a server (2000) are identical, and a second content image having a second identification difficulty based on the fact that information corresponding to one or more objects identified from the content image through a second artificial intelligence model and information corresponding to one or more objects identified from the content image through an artificial intelligence model stored in a server are identical. According to one embodiment of the present disclosure, a learning image may include a first learning image corresponding to a first content image having a first identification difficulty and a second learning image corresponding to a second content image having a second identification difficulty.

[0315] According to one embodiment of the present disclosure, by having at least one processor (120) execute instructions alone or in cooperation, the electronic device (1000) can receive from a server either a first type artificial intelligence model and a second type artificial intelligence model learned based on each of a first type dataset and a second type dataset having different ratios between a first training image and a second training image based on a first artificial intelligence model.

[0316] According to one embodiment of the present disclosure, by having at least one processor (120) execute instructions alone or in cooperation, the electronic device (1000) can transmit the test results of either a first type artificial intelligence model and a second type artificial intelligence model to a server.

[0317] According to one embodiment of the present disclosure, by having at least one processor (120) execute instructions alone or in cooperation, the electronic device (1000) can receive from a server information corresponding to a final model determined based on the test results of either the first type artificial intelligence model or the second type artificial intelligence model among the first artificial intelligence model, the first type artificial intelligence model, and the second type artificial intelligence model. According to one embodiment of the present disclosure, by having at least one processor (120) execute instructions alone or in cooperation, the electronic device (1000) can update the first artificial intelligence model based on the information corresponding to the received final model.

[0318] According to one embodiment of the present disclosure, a server (2000) is provided.

[0319] According to one embodiment of the present disclosure, a server (2000) may include a communication interface (210) that communicates with at least one electronic device, at least one processor (220) that includes a processing circuit, and a memory (230) that stores a plurality of instructions.

[0320] According to one embodiment of the present disclosure, by having at least one processor (220) execute instructions alone or in cooperation, the server (2000) can receive information regarding one or more objects identified from the content image through the content image and one or more artificial intelligence models stored in the at least one electronic device.

[0321] According to one embodiment of the present disclosure, by having at least one processor (220) execute instructions either alone or in cooperation, the server (2000) can identify one or more objects from a content image through an artificial intelligence model stored in the server (2000) and extract an incorrect answer area from the content image corresponding to an object identified differently from at least one electronic device and the server (2000).

[0322] According to one embodiment of the present disclosure, by having at least one processor (220) execute instructions either alone or in cooperation, the server (2000) can generate and store a learning image based on a content image and an incorrect answer area.

[0323] According to one embodiment of the present disclosure, by having at least one processor (220) execute instructions either alone or in cooperation, the server (2000) can learn a model corresponding to an artificial intelligence model stored in an electronic device using a stored training image.

[0324] According to one embodiment of the present disclosure, by having at least one processor (220) execute instructions alone or in cooperation, the server (2000) can generate a base image based on features extracted from the remaining area excluding the incorrect answer area of ​​the content image.

[0325] According to one embodiment of the present disclosure, by having at least one processor (220) execute instructions either alone or in cooperation, the server (2000) can generate a learning image by concretizing a base image based on features extracted from an incorrect answer area among content images.

[0326] According to one embodiment of the present disclosure, by having at least one processor (220) execute instructions alone or in cooperation, the server (2000) can generate a first feature prompt corresponding to the remaining area of ​​the content image.

[0327] According to one embodiment of the present disclosure, by having at least one processor (220) execute instructions alone or in cooperation, the server (2000) can generate a base image based on a first feature prompt.

[0328] According to one embodiment of the present disclosure, by having at least one processor (220) execute instructions alone or in cooperation, the server (2000) can generate a second feature prompt corresponding to an incorrect answer area in a content image.

[0329] According to one embodiment of the present disclosure, by having at least one processor (220) execute instructions alone or in cooperation, the server (2000) can generate a training image from a base image based on a second feature prompt.

[0330] According to one embodiment of the present disclosure, by having at least one processor (220) execute instructions either alone or in cooperation, the server (2000) can identify one or more objects from a training image through an artificial intelligence model stored in the server (2000) and store information corresponding to one or more objects identified from the training image as a response (or correct answer) to the object identification of the training image.

[0331] According to one embodiment of the present disclosure, by having at least one processor (220) execute instructions either alone or in cooperation, the server (2000) can determine the identification difficulty of the content image as a first identification difficulty if the information corresponding to one or more objects identified from the content image through a first artificial intelligence model and the information corresponding to one or more objects identified from the content image through an artificial intelligence model stored in the server (2000) are identical (or correspond to each other).

[0332] According to one embodiment of the present disclosure, by having at least one processor (220) execute instructions either alone or in cooperation, the server (2000) can determine the identification difficulty of the content image as a second identification difficulty if the information corresponding to one or more objects identified from the content image through the second artificial intelligence model and the information corresponding to one or more objects identified from the content image through the artificial intelligence model stored in the server (2000) are identical (or correspond to each other).

[0333] According to one embodiment of the present disclosure, by having at least one processor (220) execute instructions alone or in cooperation, the server (2000) can determine the identification difficulty of the learning image as the first identification difficulty if the content image corresponding to the learning image has a first identification difficulty, and determine the identification difficulty of the learning image as the second identification difficulty if the content image corresponding to the learning image has a second identification difficulty.

[0334] According to one embodiment of the present disclosure, by having at least one processor (220) execute instructions alone or in cooperation, the server (2000) can learn a model corresponding to a first artificial intelligence model based on each of a first type dataset and a second type dataset, wherein the ratio between a training image having a first identification difficulty and a training image having a second identification difficulty is different.

[0335] According to one embodiment of the present disclosure, by having at least one processor (220) execute instructions alone or in cooperation, the server (2000) can store a first type artificial intelligence model learned based on a first type dataset and a second type artificial intelligence model learned based on a second type dataset.

[0336] According to one embodiment of the present disclosure, by having at least one processor (220) execute instructions alone or in cooperation, the server (2000) can deploy a first type artificial intelligence model to some of at least one electronic device and deploy a second type artificial intelligence model to other of at least one electronic device.

[0337] According to one embodiment of the present disclosure, by having at least one processor (220) execute instructions alone or in cooperation, the server (2000) can receive information regarding the test results of a first type artificial intelligence model from some of at least one electronic device.

[0338] According to one embodiment of the present disclosure, by having at least one processor (220) execute instructions either alone or in cooperation, the server (2000) can receive information regarding the test results of a second type artificial intelligence model from another part of at least one electronic device.

[0339] According to one embodiment of the present disclosure, by having at least one processor (220) execute instructions alone or in cooperation, the server (2000) can receive information regarding the test results of a first artificial intelligence model from the remainder, excluding some and other parts of at least one electronic device.

[0340] According to one embodiment of the present disclosure, by having at least one processor (220) execute instructions alone or in cooperation, the server (2000) can determine a final model among the first artificial intelligence model, the first artificial intelligence model, and the second artificial intelligence model based on information regarding the test results of the first type artificial intelligence model, information regarding the test results of the second type artificial intelligence model, and information regarding the test results of the first artificial intelligence model.

[0341] According to one embodiment of the present disclosure, by having at least one processor (220) execute instructions alone or in cooperation, the server (2000) can distribute the determined first type artificial intelligence model or the second type artificial intelligence model to at least one electronic device if the determined final model is a first type artificial intelligence model or a second type artificial intelligence model.

[0342] According to one embodiment of the present disclosure, by having at least one processor (220) execute instructions alone or in cooperation, the server (2000) can learn a model corresponding to a second artificial intelligence model based on each of a third type dataset and a fourth type dataset, wherein the ratio between a training image having a first identification difficulty and a training image having a second identification difficulty is different.

[0343] According to one embodiment of the present disclosure, a content image received from at least one electronic device may correspond to an image in which one or more objects identified through a first artificial intelligence model and one or more objects identified through a second artificial intelligence model are different from each other.

[0344] According to one embodiment of the present disclosure, a system (100) is provided.

[0345] In one embodiment of the present disclosure, the system (100) may include at least one electronic device (1000) and a server (2000). In one embodiment of the present disclosure, each of the at least one electronic device (1000) may include a communication interface (110) that communicates with the server (2000), at least one processor (120) that includes a processing circuit, and a memory (130) that stores a plurality of instructions. In one embodiment of the present disclosure, the server (2000) may include a communication interface (210) that communicates with at least one electronic device (1000), at least one processor (220), and a memory (230) that stores a plurality of instructions.

[0346] In one embodiment of the present disclosure, each of at least one electronic device (1000) can identify one or more objects from an input content image using each of a first artificial intelligence model and a second artificial intelligence model by having at least one processor (120) execute instructions alone or in cooperation.

[0347] In one embodiment of the present disclosure, each of the at least one electronic device (1000) can transmit a content image, information corresponding to one or more objects identified through the first artificial intelligence model, and information corresponding to one or more objects identified through the second artificial intelligence model to a server (2000) when one or more objects identified through the first artificial intelligence model and one or more objects identified through the second artificial intelligence model do not correspond to each other by having at least one processor (120) execute instructions alone or in cooperation.

[0348] In one embodiment of the present disclosure, the server (2000) can identify one or more objects from a content image through a third artificial intelligence model by having at least one processor (220) execute instructions alone or in cooperation, and can extract incorrect answer areas from the content image corresponding to objects identified differently from each other in at least one electronic device (1000) and the server (2000).

[0349] In one embodiment of the present disclosure, the server (2000) can generate and store a learning image based on a content image and an incorrect answer area by having at least one processor (220) execute instructions alone or in cooperation.

[0350] In one embodiment of the present disclosure, the server (2000) can learn a model corresponding to each of the first artificial intelligence model and the second artificial intelligence model using a stored training image by having at least one processor (220) execute instructions alone or in cooperation.

[0351] According to one aspect of one embodiment of the present disclosure, a method of operating an electronic device (1000) is provided.

[0352] In one embodiment of the present disclosure, the method of operation of the electronic device (1000) may include the step (S210) of identifying one or more objects through each of a first artificial intelligence model and a second artificial intelligence model based on a content image input to the electronic device (1000). In one embodiment of the present disclosure, the method of operation of the electronic device (1000) may include the step (S220) of transmitting to a server, if the one or more objects identified through the first artificial intelligence model and the one or more objects identified through the second artificial intelligence model do not correspond to each other, the content image, information corresponding to one or more objects identified through the first artificial intelligence model, and information corresponding to one or more objects identified through the second artificial intelligence model. In one embodiment of the present disclosure, the method of operation of the electronic device (1000) may include the step (S230) of receiving from a server information corresponding to an update of at least one of the first artificial intelligence model or the second artificial intelligence model, which is obtained using a training image based on the content image, information corresponding to one or more objects identified through the first artificial intelligence model, and information corresponding to one or more objects identified through the second artificial intelligence model.

[0353] In one embodiment of the present disclosure, the method of operation of the electronic device (1000) may include the step (S610) of receiving an input corresponding to object identification for each of a plurality of content images. In one embodiment of the present disclosure, the method of operation of the electronic device (1000) may include the step (S620) of executing each of a first artificial intelligence model and a second artificial intelligence model in accordance with the reception of an input corresponding to object identification when the resolution of the plurality of content images is less than a threshold value, and executing the first artificial intelligence model in accordance with the reception of an input corresponding to object identification when the resolution of the plurality of content images is greater than or equal to a threshold value, and executing the second artificial intelligence model in accordance with a predetermined frequency.

[0354] In one embodiment of the present disclosure, the method of operation of the electronic device (1000) may include the step (S710) of receiving an input corresponding to object identification for each of the plurality of content images. In one embodiment of the present disclosure, the method of operation of the electronic device (1000) may include the step (S720) of alternately executing a first artificial intelligence model and a second artificial intelligence model upon receiving an input corresponding to object identification. In one embodiment of the present disclosure, the method of operation of the electronic device (1000) may include the step (S730) of executing both the first artificial intelligence model and the second artificial intelligence model when the frequency of execution of the first artificial intelligence model and the second artificial intelligence model corresponds to a predetermined frequency.

[0355] In one embodiment of the present disclosure, the method of operation of an electronic device (1000) may include the step (S1210) of storing information corresponding to a content image and one or more objects identified through a second artificial intelligence model when one or more objects identified through a first artificial intelligence model and one or more objects identified through a second artificial intelligence model do not correspond to each other. In one embodiment of the present disclosure, the method of operation of an electronic device (1000) may include the step (S1220) of learning a first artificial intelligence model by inputting information corresponding to one or more objects identified through a second artificial intelligence model as a response to object identification of a content image.

[0356] In one embodiment of the present disclosure, information corresponding to an update of at least one of a first artificial intelligence model and a second artificial intelligence model may include at least one of a model updated based on information corresponding to a training image and one or more objects identified from the training image via a server based on the first artificial intelligence model, or a model updated based on information corresponding to a training image and one or more objects identified from the training image via a server based on the second artificial intelligence model.

[0357] In one embodiment of the present disclosure, the learning image may be based on an incorrect answer area and a content image corresponding to an object identified differently between an electronic device and a server.

[0358] In one embodiment of the present disclosure, the learning image may correspond to an image materialized based on a second feature prompt corresponding to an incorrect answer area, from a base image based on a first feature prompt corresponding to the remaining area excluding the incorrect answer area among the content images.

[0359] In one embodiment of the present disclosure, a content image may include a first content image having a first identification difficulty based on the fact that information corresponding to one or more objects identified from the content image through a first artificial intelligence model and information corresponding to one or more objects identified from the content image through an artificial intelligence model stored in a server (2000) are identical, and a second content image having a second identification difficulty based on the fact that information corresponding to one or more objects identified from the content image through a second artificial intelligence model and information corresponding to one or more objects identified from the content image through an artificial intelligence model stored in a server (2000) are identical. In one embodiment of the present disclosure, a learning image may include a first learning image corresponding to a first content image having a first identification difficulty and a second learning image corresponding to a second content image having a second identification difficulty.

[0360] In one embodiment of the present disclosure, a method of operation of an electronic device (1000) may include receiving from a server either a first type artificial intelligence model and a second type artificial intelligence model trained based on each of a first type dataset and a second type dataset, wherein the ratio between a first training image and a second training image is different based on a first artificial intelligence model. In one embodiment of the present disclosure, a method of operation of an electronic device (1000) may include a step of transmitting a test result of either a first type artificial intelligence model and a second type artificial intelligence model to a server. In one embodiment of the present disclosure, a method of operation of an electronic device (1000) may include a step of receiving from a server information corresponding to a final model determined based on a test result of either a first type artificial intelligence model or a second type artificial intelligence model among a first artificial intelligence model, a first type artificial intelligence model, and a second type artificial intelligence model. In one embodiment of the present disclosure, a method of operation of an electronic device (1000) may include a step of updating at least one of a first artificial intelligence model or a second artificial intelligence model based on information corresponding to the final model.

[0361] According to one aspect of one embodiment of the present disclosure, a method of operating a server (2000) is provided.

[0362] In one embodiment of the present disclosure, the method of operation of a server (2000) may include the step (S310) of receiving information corresponding to one or more objects identified from a content image through a content image and one or more artificial intelligence models stored in at least one electronic device.

[0363] In one embodiment of the present disclosure, the method of operation of the server (2000) may include the step (S320) of identifying one or more objects from a content image through an artificial intelligence model stored in the server (2000) and extracting an incorrect answer area from the content image corresponding to an object identified differently from at least one electronic device and the server (2000).

[0364] In one embodiment of the present disclosure, the method of operation of the server (2000) may include the step (S330) of generating and storing a learning image based on a content image and an incorrect answer area.

[0365] In one embodiment of the present disclosure, the method of operation of the server (2000) may include the step (S340) of learning a model corresponding to one or more artificial intelligence models stored in an electronic device using a stored training image.

[0366] In one embodiment of the present disclosure, the method of operation of the server (2000) may include the step (S810) of generating a base image based on features extracted from the remaining area excluding the incorrect answer area among the content images.

[0367] In one embodiment of the present disclosure, the method of operation of the server (2000) may include the step (S820) of generating a learning image by specifying a base image based on features extracted from an incorrect answer area among content images.

[0368] In one embodiment of the present disclosure, the method of operation of a server (2000) may include the step (S910) of determining the identification difficulty of a content image as a first identification difficulty if the information corresponding to one or more objects identified from a content image through a first artificial intelligence model included in at least one electronic device and the information corresponding to one or more objects identified from a content image through an artificial intelligence model stored in the server (2000) are identical (or correspond to each other).

[0369] In one embodiment of the present disclosure, the method of operation of a server (2000) may include the step (S920) of determining the identification difficulty of a content image as a second identification difficulty if the information corresponding to one or more objects identified from a content image through a second artificial intelligence model included in at least one electronic device and the information corresponding to one or more objects identified from a content image through an artificial intelligence model stored in the server (2000) are identical (or correspond to each other).

[0370] In one embodiment of the present disclosure, the method of operation of the server (2000) may include the step (S930) of determining the identification difficulty of the learning image as the first identification difficulty if the content image corresponding to the learning image has a first identification difficulty, and determining the identification difficulty of the learning image as the second identification difficulty if the content image corresponding to the learning image has a second identification difficulty.

[0371] According to one embodiment of the present disclosure, a computer-readable recording medium is provided on which a program for performing a method of operating an electronic device (1000) is recorded.

[0372] According to one embodiment of the present disclosure, a computer-readable recording medium (e.g., a computer-readable non-transient recording medium) is provided on which a program for performing a method of operation of a server (2000) is recorded.

[0373] A device-readable storage medium (e.g., a device-readable non-transitory recording medium) may be provided in the form of a non-transitory storage medium. Here, 'non-transitory storage medium' may mean a tangible device that does not contain a signal (e.g., electromagnetic waves), and this term may not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily. For example, a 'non-transitory storage medium' may include a buffer in which data is stored temporarily.

[0374] According to one embodiment of the present disclosure, the method of operation according to the various embodiments disclosed herein may be provided as included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., downloadable app) may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

Claims

1. In an electronic device (1000), A communication interface (110) including a communication circuit; Memory (130) for storing multiple instructions; and It includes at least one processor (120) comprising a processing circuitry and operably coupled to the memory (130); and By having at least one processor (120) execute the instructions either alone or in cooperation, the electronic device (1000) Based on the content image input to the electronic device (1000), one or more objects are identified through each of the first artificial intelligence model and the second artificial intelligence model, and Based on the fact that one or more objects identified through the first artificial intelligence model and one or more objects identified through the second artificial intelligence model do not correspond to each other, the content image, information corresponding to one or more objects identified through the first artificial intelligence model, and information corresponding to one or more objects identified through the second artificial intelligence model are transmitted to a server through the communication interface (110). An electronic device (1000) that receives from the server information corresponding to an update of at least one of the first artificial intelligence model or the second artificial intelligence model, obtained using a learning image generated based on the content image, information corresponding to one or more objects identified through the first artificial intelligence model, and information corresponding to one or more objects identified through the second artificial intelligence model, through the communication interface (110).

2. In Paragraph 1, By having at least one processor (120) execute the instructions either alone or in cooperation, the electronic device (1000) Receive input corresponding to object identification for each of multiple content images, and An electronic device (1000) that executes each of the first artificial intelligence model and the second artificial intelligence model upon receiving an input corresponding to object identification based on the resolution of the plurality of content images being less than a threshold value, and executes the first artificial intelligence model upon receiving an input corresponding to object identification based on the resolution of the plurality of content images being greater than or equal to the threshold value, and executes the second artificial intelligence model according to a predetermined frequency.

3. In Paragraph 1, By having at least one processor (120) execute the instructions either alone or in cooperation, the electronic device (1000) Receive input corresponding to object identification for each of multiple content images, and The first artificial intelligence model and the second artificial intelligence model are executed alternately upon receiving an input corresponding to the object identification, and An electronic device (1000) that executes both the first artificial intelligence model and the second artificial intelligence model based on the frequency of execution of the first artificial intelligence model and the second artificial intelligence model corresponding to a preset frequency.

4. In any one of paragraphs 1 to 3, By having at least one processor (120) execute the instructions either alone or in cooperation, the electronic device (1000) If one or more objects identified through the first artificial intelligence model and one or more objects identified through the second artificial intelligence model do not correspond to each other, information corresponding to the content image and one or more objects identified through the second artificial intelligence model is stored, and An electronic device (1000) that learns the first artificial intelligence model by inputting information corresponding to one or more objects identified through the second artificial intelligence model as a response to object identification of the content image.

5. In any one of paragraphs 1 through 4, An electronic device (1000), wherein information corresponding to an update of at least one of the first artificial intelligence model and the second artificial intelligence model comprises at least one of a model updated based on the first artificial intelligence model and information corresponding to the training image and one or more objects identified from the training image through the server, or a model updated based on the second artificial intelligence model and information corresponding to the training image and one or more objects identified from the training image through the server.

6. In any one of paragraphs 1 through 5, The above learning image is an electronic device (1000) based on an incorrect answer area corresponding to an object identified differently between the electronic device (1000) and the server and the content image.

7. In Paragraph 6, The above learning image is an electronic device (1000) corresponding to an image materialized based on a second feature prompt corresponding to an incorrect answer area, from a base image based on a first feature prompt corresponding to the remaining area excluding the incorrect answer area among the content images.

8. In a method of operating an electronic device (1000), the method is: A step (S210) of identifying one or more objects through each of a first artificial intelligence model and a second artificial intelligence model based on a content image input to the electronic device (1000); Based on the fact that one or more objects identified through the first artificial intelligence model and one or more objects identified through the second artificial intelligence model do not correspond to each other, a step (S220) of transmitting to a server the content image, information corresponding to one or more objects identified through the first artificial intelligence model, and information corresponding to one or more objects identified through the second artificial intelligence model; and A method of operation of an electronic device (1000), comprising the step (S230) of receiving from the server information corresponding to an update of at least one of the first artificial intelligence model and the second artificial intelligence model, which is obtained using a training image generated based on the above content image, information corresponding to one or more objects identified through the first artificial intelligence model, and information corresponding to one or more objects identified through the second artificial intelligence model.

9. In Paragraph 8, The method of operation of the above electronic device (1000) is, A step of receiving input corresponding to object identification for each of the multiple content images (S610); and A method of operation of an electronic device (1000), further comprising the step (S620) of executing each of the first artificial intelligence model and the second artificial intelligence model upon receiving an input corresponding to object identification based on the fact that the resolution of the plurality of content images is less than a threshold value, executing the first artificial intelligence model upon receiving an input corresponding to object identification based on the fact that the resolution of the plurality of content images is greater than or equal to the threshold value, and executing the second artificial intelligence model in response to a predetermined frequency.

10. In any one of paragraphs 8 to 9, The method of operation of the above electronic device (1000) is, A step of receiving input corresponding to object identification for each of the multiple content images (S710); Step (S720) of alternately executing the first artificial intelligence model and the second artificial intelligence model upon receiving an input corresponding to the object identification; and A method of operation of an electronic device (1000), further comprising the step (S730) of executing both the first artificial intelligence model and the second artificial intelligence model based on the frequency of execution of the first artificial intelligence model and the second artificial intelligence model corresponding to a predetermined frequency.

11. In any one of paragraphs 8 through 10, The method of operation of the above electronic device (1000) is, Based on the fact that one or more objects identified through the first artificial intelligence model and one or more objects identified through the second artificial intelligence model do not correspond to each other, a step (S1210) of storing information corresponding to the content image and one or more objects identified through the second artificial intelligence model; and A method of operation of an electronic device (1000), further comprising the step (S1220) of learning the first artificial intelligence model by inputting information corresponding to one or more objects identified through the second artificial intelligence model as a response to object identification of the content image.

12. In any one of paragraphs 8 through 11, A method of operation of an electronic device (1000), wherein information corresponding to an update of at least one of the first artificial intelligence model and the second artificial intelligence model comprises at least one of a model updated based on the first artificial intelligence model and information corresponding to the training image and one or more objects identified from the training image through the server, or a model updated based on the second artificial intelligence model and information corresponding to the training image and one or more objects identified from the training image through the server.

13. In any one of paragraphs 8 through 12, The above learning image is a method of operation of the electronic device (1000) based on an incorrect answer area corresponding to an object differently identified between the electronic device (1000) and the server and the content image.

14. In Paragraph 13, A method of operation of an electronic device (1000), wherein the above learning image corresponds to an image materialized based on a second feature prompt corresponding to the incorrect answer area, from a base image based on a first feature prompt corresponding to the remaining area excluding the incorrect answer area among the above content images.

15. A computer-readable recording medium having a program for executing the method of any one of paragraphs 8 through 14 on a computer.