Method and electronic device for neural symbol learning of artificial intelligence model
By combining the update mechanism of neural loss and symbol loss in the AI model, the problem that traditional neural AI systems cannot understand visual and language concepts is solved, and efficient deployment and understanding capabilities on embedded devices are achieved.
Patent Information
- Application Number
- CN202380080360.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-19
- Filing Date
- 2023-12-18
- Publication Date
- 2025-06-27
AI Technical Summary
Traditional neural AI systems are unable to understand visual concepts and language concepts in natural language, resulting in difficulties in implementing neural symbolic artificial intelligence (NSAI) models on resource-limited devices.
The neural loss of the AI model is determined by comparing the predicted probability for the content of the input data with the predefined expected probability, and the symbol loss is determined by comparing the predicted probability with the predetermined undesired probability, thereby updating the weights of multiple layers of the AI model.
It realizes efficient deployment of AI models on embedded devices, improves the ability to understand visual concepts and language concepts, reduces the complexity of AI models, reduces power and memory consumption, and enhances the user experience.
Smart Images

Figure CN120226016A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to artificial intelligence (AI). More specifically, the present disclosure relates to the neuro-symbolic learning or training of an AI model, and to facilitating the deployment of the AI model on an embedded device. Background Art
[0002] AI models have had a profound impact on all aspects of our lives, heralding a new era of innovation. These models are being widely applied in various applications, and their optimization for edge computing and deployment on embedded devices contribute to driving their continued growth and success.
[0003] Despite significant progress in AI technology, traditional neural networks or systems still cannot understand visual concepts. For example, as Figure 1A shown, even when an object moves from one place to another, traditional neural AI systems tend to label the object as a tree. This is mainly because such systems cannot fully grasp the context information from the surrounding scenes depicted in images or videos.
[0004] In addition, traditional neural AI systems exhibit deficiencies in understanding the linguistic concepts inherent in natural language. This is illustrated in Figure 1C and Figure 1D where traditional neural AI systems cannot distinguish between "standing next to an animal" and "being chased by an animal". This shortcoming stems from the inability of traditional neural AI systems to understand the nuances of linguistic concepts.
[0005] Therefore, it is desirable to provide a mechanism for an AI model that does not have the above problems.
[0006] The above information is presented only as background information to assist in understanding the present disclosure. No determination has been made, and no assertion is made, as to whether any of the above constitutes prior art with respect to the present disclosure. Summary of the Invention
[0007] Aspects of the present disclosure are to at least address the above problems and / or disadvantages, and to at least provide the advantages described below. Accordingly, one aspect of the present disclosure is to provide neuro-symbolic learning or training of an AI model, and to facilitate the deployment of the AI model on an embedded device.
[0008] Another aspect of the present disclosure is to determine the neural loss of an AI model by comparing the predicted probability for the input data content with a predefined desired probability.
[0009] Another aspect of the present disclosure is to determine the symbolic loss of an AI model by comparing the predicted probability for the input data content with a predefined undesired probability.
[0010] Another aspect of the present disclosure is to determine the weights of multiple layers of an AI model and update the weights of the layers of the AI model based on neural loss and symbolic loss.
[0011] Additional aspects will be set forth in part in the following description, and in part will be obvious from the description, or may be learned by practice of the presented embodiments.
[0012] According to one aspect of the present disclosure, a method for neuro-symbolic learning of an artificial intelligence (AI) model is provided. The method includes receiving, by an electronic device, input data including multiple contents for neuro-symbolic learning of the AI model, determining a prediction probability for each of the multiple contents in the output of the AI model for the input data, determining a neural loss of the AI model by comparing the prediction probability for each of the multiple contents with a predefined desired probability for each of the multiple contents, determining a symbolic loss of the AI model by comparing the prediction probability for each of the multiple contents with a predefined undesired probability for each of the multiple contents, determining the weights of multiple layers of the AI model, and updating the weights of the multiple layers of the AI model based on the neural loss and the symbolic loss.
[0013] In an embodiment, the method includes determining a training loss of the AI model as a measure of the neural loss and the symbolic loss, and proportionally updating the weights of the multiple layers of the AI model according to the determined training loss.
[0014] In an embodiment, the method includes selecting an external symbolic knowledge graph including multiple common-sense facts, constructing a negative knowledge graph including multiple facts violating common sense from the external symbolic knowledge graph, identifying a set of maximum violated facts from the multiple facts violating common sense, and generating symbolic labels in the form of probabilities for the set of maximum violated facts. The set of maximum violated facts includes undesired probabilities.
[0015] In an embodiment, the multiple common-sense facts are at least one of a predefined rule set and a set of common-sense facts associated with the real world.
[0016] In an embodiment, the set of maximum violated facts includes facts contrary to multiple common-sense facts about the input data.
[0017] In an embodiment, the method includes determining multiple contents and the relationships between the multiple contents, and determining at least one scene graph based on the multiple contents and the relationships between the multiple contents. The at least one scene graph is a structural representation of the multiple contents and the relationships between the multiple contents in the input data.
[0018] In an embodiment, neural AI is used to determine a plurality of contents, and symbolic AI is used to determine the relationships between the plurality of contents, wherein the symbolic AI follows real-world common sense through an external symbolic knowledge graph.
[0019] According to another aspect of the present disclosure, an electronic device for neuro-symbolic learning of an artificial intelligence (AI) model is provided. The electronic device includes a processor, a communicator, a neuro-symbolic AI controller, and a memory that stores one or more programs including computer-executable instructions. When the computer-executable instructions are run by the processor, the electronic device is caused to perform the following operations: receiving input data including a plurality of contents for neuro-symbolic learning of an AI model, determining a prediction probability for each of the plurality of contents in the input data in the output of the AI model, determining a neural loss of the AI model by comparing the prediction probability for each of the plurality of contents with a predefined desired probability for each of the plurality of contents, determining a symbolic loss of the AI model by comparing the prediction probability for each of the plurality of contents with a predefined undesired probability for each of the plurality of contents, determining a training loss of the AI model as a measure of the neural loss and the symbolic loss, determining weights of a plurality of layers of the AI model, and updating the weights of the plurality of layers of the AI model based on the neural loss and the symbolic loss.
[0020] According to another aspect of the present disclosure, one or more non-transitory computer-readable storage media are provided, which store one or more computer programs including computer-executable instructions. When the computer-executable instructions are run by one or more processors of an electronic device, the electronic device is caused to perform operations. The operations include: receiving input data including a plurality of contents for neuro-symbolic learning of an artificial intelligence (AI) model, determining a prediction probability for each of the plurality of contents in the input data in the output of the AI model, determining a neural loss of the AI model by comparing the prediction probability for each of the plurality of contents with a predefined desired probability for each of the plurality of contents, determining a symbolic loss of the AI model by comparing the prediction probability for each of the plurality of contents with a predefined undesired probability for each of the plurality of contents, determining weights of a plurality of layers of the AI model, and updating the weights of the plurality of layers of the AI model based on the neural loss and the symbolic loss.
[0021] The description of various embodiments of the present disclosure is disclosed below in conjunction with the accompanying drawings. Other aspects, advantages, and significant features of the present disclosure will become apparent to those skilled in the art. Description of the Drawings
[0022] The above and other aspects, features, and advantages of specific embodiments of the present disclosure will become more apparent from the following description in conjunction with the accompanying drawings, wherein: Figure 1A is an example showing object classification according to the related art; Figure 1B is a schematic diagram showing a knowledge graph according to the related art; Figure 1C and Figure 1D is an example showing language concepts according to the related art; Figure 2 is a schematic diagram clarifying the comparison between traditional systems according to the related art; Figure 3 is a block diagram of an electronic device for training an AI model according to an embodiment of the present disclosure; Figure 4 is a flowchart showing a method for training an AI model according to an embodiment of the present disclosure; Figure 5 is an example showing scene graph generation for understanding visual concepts according to an embodiment of the present disclosure; Figure 6 is an example showing scene graph generation through a neuro-symbolic AI controller according to an embodiment of the present disclosure; Figure 7 is a flowchart showing a method for providing neuro-symbolic training of AI on an edge device according to an embodiment of the present disclosure; Figure 8 is a schematic diagram showing the conversion of symbolic AI into labels according to an embodiment of the present disclosure; Figure 9 is a schematic diagram showing obtaining an undesired label according to an embodiment of the present disclosure; Figure 10 is a schematic diagram showing the training process of a neuro-symbolic AI controller according to an embodiment of the present disclosure; and Figure 11 is a flowchart showing a deployment process according to an embodiment of the present disclosure.
[0023] Throughout the drawings, the same reference numerals are used to denote the same elements. Detailed Description
[0024] The following description with reference to the accompanying drawings helps to fully understand various embodiments of the present disclosure defined by the claims and their equivalents. It includes various specific details to assist in understanding, but these details are only considered exemplary. Thus, those of ordinary skill in the art will recognize that various changes and modifications can be made to the various embodiments described herein without departing from the scope and spirit of the present disclosure. In addition, descriptions of well-known functions and configurations may be omitted for clarity and conciseness.
[0025] The terms and words used in the following description and claims are not limited to their written meanings, but are used solely by the inventors to enable a clear and consistent understanding of the present disclosure. Thus, it will be apparent to those skilled in the art that the following description of the various embodiments of the present disclosure is for illustrative purposes only and is not for the purpose of limiting the present disclosure as defined by the appended claims and their equivalents.
[0026] It should be understood that, unless the context clearly dictates otherwise, the singular forms include plural referents. Thus, for example, a reference to "component surface" includes a reference to one or more such surfaces.
[0027] As is traditional in the art, embodiments may be described and illustrated with blocks that perform the described functions. These blocks, which may be referred to herein as units or modules, etc., are physically implemented by analog or digital circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuits, etc., and may optionally be driven by firmware. The circuits may be embodied, for example, in one or more semiconductor chips or on a substrate support such as a printed circuit board, etc. The circuits constituting the blocks may be implemented by dedicated hardware, or by a processor (e.g., one or more programmed microprocessors and associated circuits), or by a combination of dedicated hardware performing some of the functions of the blocks and a processor performing other functions of the blocks. Without departing from the scope of the present disclosure, each block of an embodiment may be physically divided into two or more interacting and discrete blocks. Similarly, without departing from the scope of the present disclosure, the blocks of an embodiment may be physically combined into more complex blocks.
[0028] The drawings are used to assist in an easy understanding of the various technical features, and it should be understood that the embodiments presented herein are not limited by the drawings. Thus, the present disclosure should be construed as extending to any variations, equivalents, and alternatives in addition to the embodiments specifically set forth in the drawings. Although terms such as first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are generally only used to distinguish one element from another.
[0029] Throughout the disclosure, the term "neural loss" refers to the value of any general standard loss metric obtained by comparing the output (prediction) of a neural network with the desired output or ground truth provided as part of the dataset used for training.
[0030] Throughout the disclosure, the term "symbolic loss" refers to the value of any general standard loss metric obtained by comparing the output (prediction) of a neural network with an undesired output or symbolic label obtained from an external knowledge graph using neural-guided projection.
[0031] Throughout the disclosure, the term "training loss" refers to the aggregation of the neural loss and the symbolic loss. The general aggregation rule can be to add both the neural loss and the symbolic loss, but it doesn't have to be specifically this addition. It can be any possible form of aggregation.
[0032] Accordingly, embodiments herein disclose a method for neuro-symbolic learning or training an AI model and enabling the deployment of the AI model on an embedded device. The method includes an electronic device receiving input data including various contents for training the AI model. Additionally, the method includes the electronic device determining a prediction probability for each content in the input data's contents in the output of the AI model. Additionally, the method includes the electronic device determining the neural loss of the AI model by comparing the prediction probability for each content with a predefined desired probability for each content. Additionally, the method includes the electronic device determining the symbolic loss of the AI model by comparing the prediction probability for each content with a predefined undesired probability for each content. Additionally, the method includes the electronic device determining the training loss of the AI model as a measure of the neural loss and the symbolic loss. Additionally, the method includes the electronic device determining the weights of one or more layers of the AI model. Additionally, the method includes the electronic device proportionally updating the weights of one or more layers of the AI model according to the determined training loss.
[0033] Accordingly, embodiments herein disclose an electronic device for training an AI model. The electronic device includes a memory, a processor coupled to the memory, and a communicator coupled to the memory and the processor. The electronic device includes a neuro-symbolic AI controller coupled to the memory, the processor, and the communicator. The neuro-symbolic AI controller is configured to receive input data including various contents for neuro-symbolic learning or training an AI model. Additionally, the neuro-symbolic AI controller determines a prediction probability for each content in the input data's contents in the output of the AI model. Additionally, the neuro-symbolic AI controller determines the neural loss of the AI model by comparing the prediction probability for each content with a predefined desired probability for each content. Additionally, the neuro-symbolic AI controller determines the symbolic loss of the AI model by comparing the prediction probability for each content with a predefined undesired probability for each content. Additionally, the neuro-symbolic AI controller determines the training loss of the AI model as a measure of the neural loss and the symbolic loss. Additionally, the neuro-symbolic AI controller determines the weights of one or more layers of the AI model and proportionally updates the weights of the layers of the AI model according to the determined training loss.
[0034] Traditional techniques and systems for training neural networks lack provisions for implementing AI model training. Additionally, such traditional methods do not address the problem of implementing a neuro-symbolic artificial intelligence (NSAI) model on resource-constrained devices.
[0035] Traditional methods and systems cannot meet on-device learning because they lack the ability to understand techniques and understand visual and language contexts.
[0036] Compared with traditional techniques and systems, the proposed disclosure incorporates a symbolic AI method into the training process.
[0037] The proposed disclosure differs from traditional methods and systems by constructing a neuro-symbolic AI model that needs to be trained using a symbolic knowledge graph. The proposed disclosure provides an efficient on-device learning strategy.
[0038] Compared with traditional methods and systems, the proposed disclosure can create NSAI models that do not require a symbolic knowledge graph for inference, thus facilitating their deployment on embedded devices. The proposed method integrates a commonsense knowledge base to improve the accuracy and efficiency of neural AI.
[0039] Different from traditional methods and systems, the proposed disclosure promotes proficient on-device learning by leveraging neuro-symbolic techniques. Notably, the method uses a neuro-symbolic approach to construct a scene graph, thereby enabling improved understanding of input images.
[0040] Compared with traditional methods and systems, the proposed disclosure includes training data as well as a symbolic knowledge base to train the neuro-symbolic method. Optionally, the proposed method constructs a scene graph as a structural representation of the input data, making it different from traditional methods and systems. A further difference is that the proposed method achieves comparable accuracy to traditional systems while requiring less data. The reduction in required data is particularly important for on-device learning and user personalization, where user-generated data is typically limited in quantity.
[0041] In traditional methods, a neural AI model is required to learn connection relationships and symbolic tasks, resulting in a larger and more complex model. However, our proposed method differs from traditional methods by implementing symbolic task learning based on an external knowledge graph. Therefore, the proposed disclosure effectively reduces the complexity of the AI model, thereby reducing power consumption and memory consumption, enhancing latency, and improving the user experience.
[0042] Compared with traditional technologies and systems, the proposed method includes training using an external knowledge graph. Optionally, it belongs to the field of explainable AI, providing a new perspective for building and deploying NSAI models on embedded devices. In addition, the proposed method allows for the convenient deployment and inference of neuro-symbolic AI models on resource-constrained edge devices and embedded devices such as smartphones. The method is capable of building simpler neural AI models that are regulated with symbolic knowledge, resulting in explainable results. Moreover, the technology can build AI models in data-scarce environments, thus promoting effective personalization. By incorporating symbolic knowledge to design an explainable, accurate, lightweight, and user-personalized neuro-symbolic AI for edge devices, the proposed method differentiates itself from traditional methods. Finally, it allows for effective personalization and on-device learning in data-scarce environments.
[0043] It should be understood that each block in the flowcharts and combinations of flowcharts can be executed by one or more computer programs including instructions. All of the one or more computer programs can be stored in a single memory, or the one or more computer programs can be divided into different parts and stored in multiple different memories.
[0044] Any function or operation described herein can be processed by one processor or a combination of processors. One processor or a combination of processors is a circuit that performs processing and includes circuits such as: an application processor (AP, e.g., a central processing unit (CPU)), a communication processor (CP, e.g., a modem), a graphics processing unit (GPU), a neural processing unit (NPU) (e.g., an artificial intelligence (AI) chip), a Wi-Fi chip, a Bluetooth TM chip, a global positioning system (GPS) chip, a near-field communication (NFC) chip, a connection chip, a sensor controller, a touch controller, a fingerprint sensor controller, a display driver integrated circuit (IC), an audio codec chip, a universal serial bus (USB) controller, a camera controller, an image processing IC, a microprocessor unit (MPU), a system-on-chip (SoC), an integrated circuit (IC), etc.
[0045] Figure 1A is an example showing object classification according to the related art.
[0046] Refer to Figure 1A , even when person 11 is in motion, traditional AI wrongly classifies person 11 as a tree. This is due to the limited ability of traditional systems to understand visual abstractions. Unfortunately, the AI model has never absorbed the basic common-sense principle that a tree does not have the ability to walk.
[0047] Figure 1B is a schematic diagram showing the knowledge graph according to the related art.
[0048] Referring to Figure 1B , the domain of the knowledge graph contains a large amount of factual and common-sense information. Symbolic AI, typically measured in gigabytes (GB), relies on these knowledge graphs to incorporate common sense. However, it is impossible for NSAI to deploy these knowledge graphs on the device.
[0049] Figure 1C and Figure 1D are examples showing language concepts according to the related art.
[0050] Referring to Figure 1C and Figure 1D , traditional systems cannot understand language concepts. For example, traditional systems cannot distinguish between Figure 1C "standing next to an animal" in Figure 1D and
[0051] "being chased by an animal" in
[0052] Figure 2 This is because traditional systems cannot understand language concepts. Traditional neural AI also cannot distinguish context in language because the AI model has never learned the difference between "standing next to" and "chasing".
[0053] Referring to Figure 2 , during inference on the device, at 201, the traditional system 200 receives an image. At 202, the neuro-symbolic model is presented with the image for generating a scene graph. At 203, the image is sent to the cloud by the traditional system. At 204, the traditional system accesses an external knowledge graph. At 205, the traditional system generates a scene graph via the external knowledge graph. At 206, the traditional system obtains the scene graph from the cloud and sends it to the embedded device. Subsequently, at 207, the traditional system operates on the scene graph according to the application requirements. At 208, the traditional system generates an accurate scene graph with a large latency and sends it to the embedded device.
[0054] During inference on the device, at 209, the proposed system 223 receives an image. At 210, the pure neural AI system incorporates symbolic information into its parameters. At 211, the resulting scene graph is then generated on the device. Subsequently at 212, the proposed system operates on the scene graph according to the application requirements. At 213, finally, the proposed system generates an accurate scene graph with minimal latency.
[0055] In addition, a neural-guided projection 214 is described. At 215, the knowledge graph includes a set of common-sense facts (e.g., considering a man wearing a suit). At 216, a negative knowledge graph including facts that completely violate common sense is obtained (i.e., an illogical set of facts is obtained (e.g., a suit wearing a man)). At 217, the most illogical fact is generated (e.g., a man wearing a woman). At 218, the system generates symbolic labels that are undesirable and used during the training of the AI model.
[0056] The proposed method generates context-aware interpretable AI 219 on edge devices. In addition, it provides an AI model 220 that is both accurate and fast but not complex. Additionally, it provides an energy- and memory-efficient AI model 221 and a data-efficient AI model 222.
[0057] Generally, when compared with traditional neural AI, the technical impact of neural-guided projection (NGP) is more precise (e.g., the accuracy increases by up to 30%). Additionally, when compared with traditional NSAI methods where the NSAI method cannot be deployed on the device, the proposed system exceeds the traditional NSAI method in terms of accuracy (e.g., the accuracy exceeds 21%). Furthermore, even in the case of data reduction (e.g., data reduction by 50%), the accuracy degradation of NGP in the proposed system is better than that of traditional neural AI (e.g., about 3X times).
[0058] The proposed system provides multiple benefits, including enhanced accuracy, performance, and power efficiency, while also addressing issues such as memory latency. These advantages can be applied to various use cases in computer vision and natural language processing (NLP). For example, the system is capable of implementing advanced functions such as image library search, autonomous vehicles, robots, and visual aids for the visually impaired. These functions could not be achieved previously with traditional neural AI models and other symbolic methods. In addition, incorporating the NSAI model on the device will not only enhance the user experience but also facilitate on-device learning and personalization, even in the case of limited data.
[0059] Now referring to the drawings, and more specifically to Figures 3 to 11 , in which like reference numerals consistently denote corresponding features throughout the drawings, a preferred embodiment is shown.
[0060] Figure 3 is a block diagram of an electronic device for training an AI model according to an embodiment of the present disclosure.
[0061] Referring to Figure 3, the electronic device 300 includes a memory 301, a processor 303, a communicator 302, and a neuro-symbolic AI controller 304. The neuro-symbolic AI controller 304 is implemented by processing circuitry (such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuits, etc.), and can optionally be driven by firmware. The circuitry can be embodied, for example, in one or more semiconductors.
[0062] The memory 301 stores instructions to be run by the processor 303. The memory 301 can include non-volatile storage elements. Examples of such non-volatile storage elements can include magnetic hard disks, optical disks, floppy disks, flash memory, or in the form of electrically programmable memory (EPROM) or electrically erasable programmable (EEPROM) memory. Additionally, in some examples, the memory 301 can be considered a non-transitory storage medium. The term "non-transitory" can indicate that the storage medium is not embodied in a carrier wave or a propagated signal. However, the term "non-transitory" should not be construed to mean that the memory 301 is immovable. In some examples, the memory 301 stores a larger amount of information. In a particular example, the non-transitory storage medium can store data that can change over time (e.g., in random access memory (RAM) or a cache).
[0063] The processor 303 communicates with the memory 301, the communicator 302, and the neuro-symbolic AI controller 304. The processor 303 runs the instructions stored in the memory 301 and performs various processes. The processor 303 can include one or more processors, which can be general-purpose processors (such as a central processing unit (CPU), an application processor (AP), and similar processors) and graphics processing units only (such as a graphics processing unit (GPU), a vision processing unit (VPU)) or artificial intelligence (AI) dedicated processors (such as a neural processing unit (NPU)).
[0064] The communicator 302 includes electronic circuitry dedicated to implementing standards for wired or wireless communication. The communicator 302 performs internal communication between the internal hardware components of the electronic device 300 and external devices through one or more networks.
[0065] In an embodiment, the neuro-symbolic AI controller 304 includes a receiver 305, a probability determiner 306, a training loss determiner 307, and a weight updater 308.
[0066] The receiver 305 receives input data including various contents for training or learning an AI model. The probability determiner 306 determines the prediction probability for each content in the contents of the input data in the output of the AI model. The training loss determiner 307 determines the neural loss of the AI model by comparing the prediction probability for each content with a predetermined desired probability, and also evaluates the symbolic loss by comparing the prediction probability for each content with a predetermined undesired probability. In addition, the training loss determiner 307 determines the training loss of the AI model as a measure of the neural loss and the symbolic loss. The training loss serves as an indicator for both the neural loss and the symbolic loss. The weight updater 308 determines the weights of multiple layers of the AI model and updates the weights of one or more layers of the AI model proportionally according to the determined training loss.
[0067] The neuro-symbolic AI controller 304 selects an external symbolic knowledge graph including common-sense facts. In addition, the neuro-symbolic AI controller 304 constructs a negative knowledge graph including violations of common-sense facts from the external symbolic knowledge graph. In addition, the neuro-symbolic AI controller 304 identifies a set of maximum violations from the violations of common-sense facts. In addition, the neuro-symbolic AI controller 304 generates symbolic labels for the set of maximum violations, where the maximum violations are undesired and where the symbolic labels are provided in the form of probabilities.
[0068] In an embodiment, the electronic device 300 receives knowledge graphs including personal user information from a user. Then, these knowledge graphs are used to train the electronic device 300 to enhance its functions.
[0069] In an embodiment, the trained AI model is implemented on an edge device or a server. The input data in the embodiment contains various forms, such as but not limited to images, audio, video, and text.
[0070] In an embodiment, the common-sense facts include a predefined set of rules and a set of common-sense facts associated with the real world.
[0071] In an embodiment, the set of maximum violations includes facts that are contrary to the common-sense facts regarding the input data.
[0072] The neuro-symbolic AI controller 304 determines various contents and the relationships between various contents. Further, the neuro-symbolic AI controller 304 determines a scene graph based on the determined contents and the relationships between the contents. The scene graph is a structural representation of the various contents in the input data and the relationships between the contents.
[0073] It should be noted that the recognition of the contents is performed by the neural AI, while the discrimination of the relationships between such contents is performed by the symbolic AI following the common-sense principles of the real world. This is achieved by using an external symbolic knowledge graph.
[0074] The neuro-symbolic AI controller 304 can adopt an AI model to execute at least one of various available modules / components. The functions associated with the AI model can be executed by the memory 301 and the processor 303. A single or multiple processors supervise the processing of input data, where the processing of the input data follows a predetermined operation protocol or AI model retained in the non-volatile memory and the volatile memory. The predefined operation rules or artificial intelligence models are provided through training or learning.
[0075] Here, learning provides a predefined operation rule or AI model that represents forming desired characteristics by applying a learning process to multiple learning data. The learning can be executed in the device itself that executes the AI according to the embodiment, and / or can be implemented through a separate server / system.
[0076] The AI model can be composed of multiple neural network layers. Each layer has multiple weight values, and layer operations are performed through the calculations of the previous layer and the operations of multiple weights. Examples of neural networks include, but are not limited to, convolutional neural network (CNN), deep neural network (DNN), recurrent neural network (RNN), restricted Boltzmann machine (RBM), deep belief network (DBN), bidirectional recurrent deep neural network (BRDNN), generative adversarial network (GAN), and deep Q-network.
[0077] The learning process is a method for using multiple learning data to train a predetermined target device (e.g., a robot) to enable, permit, or control the target device to make determinations or predictions. Examples of the learning process include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0078] Although Figure 3 a series of hardware components included in the electronic device 300 are described, it should be noted that alternative embodiments are not limited to this configuration. In other embodiments, the number of components included in the electronic device 300 can vary. In addition, the labels or names assigned to each component are for illustrative purposes only and do not limit the scope of the present disclosure. Two or more components can be combined to provide the same or almost the same function in the electronic device 300.
[0079] Figure 4 is a flowchart showing a method for training an AI model according to an embodiment of the present disclosure.
[0080] Referring to Figure 4 , flowchart 400 shows that at operation 401, the electronic device 300 receives input data including various contents for neuro-symbolic learning or training the AI model. For example, in the electronic device 300 described in Figure 4 , the neuro-symbolic AI controller 304 receives an input image including multiple objects for neuro-symbolic learning or training the AI model.
[0081] At operation 402, the electronic device 300 determines the prediction probability for each content in the content of the output of the AI model for the input data. In the Figure 4 described electronic device 300, the neuro-symbolic AI controller 304 determines the expected probability for each content in the input image in the output of the AI model.
[0082] At operation 403, the electronic device 300 determines the neural loss of the AI model by comparing the prediction probability for each content with the predefined desired probability for each content. For example, the neuro-symbolic AI controller 304 determines the neural loss of the AI model based on the comparison between the prediction probability for each content in the content of the input image and the predefined probability for each content.
[0083] At operation 404, the electronic device 300 determines the symbolic loss of the AI model by comparing the prediction probability for each content with the predefined undesired probability for each content. For example, in the Figure 4 described electronic device 300, the neuro-symbolic AI controller 304 is configured to determine the symbolic loss of the AI model based on the comparison between the prediction probability for each content in the input image and the predefined undesired probability for each content.
[0084] At operation 405, the electronic device 300 determines the training loss of the AI model as a measure of the neural loss and the symbolic loss. In one embodiment, the training loss is established by applying various functions (such as addition, subtraction, multiplication, and division) to both the neural loss and the symbolic loss.
[0085] At operation 406, the electronic device 300 determines the weights of one or more layers of the AI model and updates the weights of one or more layers of the AI model proportionally according to the determined training loss. For example, in the Figure 3 described electronic device 300, the neuro-symbolic AI controller 304 determines the weights of one or more layers of the AI model proportionally according to the determined training loss.
[0086] The various actions, acts, blocks, operations, etc. in the flowchart 400 may be performed in the presented order, in a different order, or simultaneously. Additionally, in some embodiments, some of the actions, acts, blocks, operations, etc. may be omitted, added, modified, skipped, etc. without departing from the scope of the present disclosure.
[0087] Figure 5 is an example of scene graph generation for understanding visual concepts according to an embodiment of the present disclosure.
[0088] Refer to Figure 5, the generation of the scene graph is for the purpose of understanding visual concepts. In this case, the proposed method endeavors to construct a structural representation of the different elements in the image 501 by creating the scene graph 502. This is achieved by assigning the following tasks to the AI model: detecting all the objects present in the image and identifying the relationships that exist between the objects.
[0089] In an embodiment, neural AI can be utilized for object detection, and symbolic AI can be utilized for relationship detection.
[0090] Referring to Figure 5 , "a man wearing a suit" can be extracted from the scene graph as a fact. Similarly, "the man is next to the woman holding a dog" is another fact used to construct the scene graph 502.
[0091] Figure 6 is an example showing the generation of a scene graph by a neuro-symbolic AI controller according to an embodiment of the present disclosure.
[0092] Referring to Figure 6 , in operation 601, the image 501 is provided to the neural AI. For readability in the image, the DNN is simplified. In operation 602, the neuro-symbolic AI controller 304 assigns probabilities for the prediction of objects. In operation 603, the symbolic AI determines the relationships between the objects detected by the neural AI of the neuro-symbolic AI controller 304.
[0093] In subsequent operation 604, facts are determined from the symbolic AI, while in operation 605, a set of facts is derived from the same AI system to generate the scene graph. In Figure 6 the traditional neuro-symbolic training process is indicated. However, the difficulty lies in the need for symbolic AI in the training and inference processes, while it is not feasible to deploy symbolic AI on the device.
[0094] Figure 7 is a flowchart showing a proposed method for providing neuro-symbolic AI on an embedded device according to an embodiment of the present disclosure.
[0095] Referring to Figure 7 , starting at operation 701, the initial iteration is started by adopting a database including an image set and its corresponding labels of interest. Continuing to operation 702, the output of the deep neural network (DNN) is continuously evaluated. Gradually progressing to operation 703, a comparative analysis is performed between the evaluated output and the desired label, thereby deriving a neural loss.
[0096] In operation 704, the system compares the evaluated output with the undesired labels to obtain a symbolic loss, and then in operation 705, the symbolic loss and the neural loss are aggregated to obtain a training loss. The proposed neural-guided projection of the system is used to solve the problem of generating undesired labels, providing the following advantages: No knowledge graph is required during training.
[0097] Makes training faster.
[0098] Supports on-device learning.
[0099] Contains a symbolic knowledge graph.
[0100] Additional loss value adjusts the DNN.
[0101] Symbolic loss minimizes the DNN complexity.
[0102] A simpler DNN ensures better performance At operation 706, the system determines whether an update to the weights of the DNN is required. If so, the next iteration begins. Otherwise, at operation 707, the proposed system terminates the process and deploys the DNN. The challenges of deploying the DNN are addressed through a deployment process to eliminate the need for a symbolic knowledge graph, providing the following advantages: No knowledge graph is required during inference on the device.
[0103] Makes inference faster.
[0104] Enhances the user experience.
[0105] Interpretable and accurate results.
[0106] Reduces memory usage and latency.
[0107] Figure 8 Is a schematic diagram showing the conversion of symbolic AI to tags according to an embodiment of the present disclosure.
[0108] Refer to Figure 8 , at operation 801, let's consider that there are four objects (i.e., man, woman, dog, and suit) present in the knowledge graph. Additionally, at operation 803, a similar approach should be adopted for all the relationships (i.e., wearing, holding, and next to) present in the knowledge graph. At operation 803, three relationships are adopted as 3-element vectors, where all positions are zero except for one position corresponding to one relationship.
[0109] Using this representation method, each fact can be represented as a probability vector 802, which is represented as a tag. These tags for the facts are obtained by combining various tags (such as the woman holding the dog 805 and the man wearing the suit 804). Conversely, at operation 806, the vector can also be decomposed to obtain all possible permutations of the facts, regardless of whether the facts are reasonable or meaningless. At operation 807, as an example of a meaningful combination, "the man holding the dog" is provided. At operation 808, an example of a meaningless combination "the man wearing the woman" is provided.
[0110] Reference Figure 8 ,the symbolic knowledge graph can be converted into a probability vector called a tag. It is also recognized that combinations of vectors can result in both meaningful and meaningless facts. Thus, starting from the symbolic knowledge graph, a large number of meaningless facts can be prepared.
[0111] Consider Figure 8 the instance where four objects interact through three relationships, resulting in a total of 64 facts, some of which are significant and some of which are not. Subsequently, after extracting the meaningful facts present in the knowledge graph, the remaining set of facts that contradict common sense is identified as integrity constraints. These constraints indicate combinations that have no links in the knowledge graph.
[0112] Figure 9 is a schematic diagram showing the obtaining of an undesired tag according to an embodiment of the present disclosure.
[0113] Reference Figure 9 ,in an embodiment, at operation 901, the requirement is to predict the facts in the image rather than all possible valid facts. For example, "a man holding a shirt" is a valid fact, but it does not exist in the image, so the AI does not need to predict it. To solve the problem in the proposed method, during training, the given information includes the image and its facts. The desired facts include all valid and meaningful facts included in the given dataset. Thus, "a man holding a shirt" is not included here. The desired facts are "a man wearing a suit", "a woman holding a dog", and "a man next to a woman". Thus, "a man holding a shirt" is not included as a desired tag. Thus, the AI will not predict "a man holding a shirt". By NGP, as Figure 9 described to create undesired facts, where the undesired facts are always invalid or meaningless. Thus, the AI model also does not predict these meaningless facts.
[0114] At operation 902, the proposed electronic device 300 analyzes a large number of integrity constraints (ICs) derived from the knowledge graph and contrasts each IC with the provided image facts. Continuing to operation 903, the electronic device 300 then calculates (such as absolute difference and addition) the loss for each IC using the knowledge graph. Among all possible combinations, the electronic device 300 selects the combination with the maximum loss, which consists of the unconnected parts in the image and the knowledge graph. An example of such a combination is "a man wearing a woman's clothes". At operation 904, the electronic device 300 endeavors to identify the meaningless combination with the highest loss from the symbolic AI. Finally, at operation 905, the electronic device 300 labels this combination as an undesired tag (i.e., a man wearing a woman's clothes).
[0115] Figure 10 is a schematic diagram showing the training process of a neuro-symbolic AI controller according to an embodiment of the present disclosure.
[0116] Refer to Figure 10 In stage 1, at operation 1011, the commonsense-based symbolic AI knowledge graph is converted into symbolic labels. The set of commonsense facts can be, for example, "A man wears a suit".
[0117] At operation 1012, integrity constraints (facts violating commonsense) or negative commonsense fact sets are identified. For example, "A suit wears a man". At operation 1013, the maximum violating fact is determined, for example, "A man wearing a woman's clothing". At operation 1014, the "symbolic" labels in the form of undesired probabilities are determined.
[0118] In stage 2, training data with images and desired labels is first used. At operation 1003, the DNN 1002 is used to determine various objects and relationships in the image 1001. At operation 1004, the prediction probability is determined, and the neural loss 1007 is calculated based on the difference between the prediction probability and the desired probability. At operation 1006, the proposed system determines symbolic labels based on the undesired probability and determines the symbolic loss 1008 of the AI model by comparing the prediction probability for each content with the undesired probability for each content. In addition, the symbolic loss 1008 is projected onto the training loss 1009, where the training loss 1009 of the AI model is a measure of the neural loss 1007 and the symbolic loss 1008. In addition, at operation 1010, the proposed system determines the weights of multiple layers of the AI model and proportionally updates the weights of one or more layers of the AI model according to the determined training loss 1009.
[0119] Figure 11 is a flowchart showing the deployment process according to an embodiment of the present disclosure.
[0120] Refer to Figure 11 At operation 1101, a model is created by combining a neural model for object detection, a CNN model for relation classification, and an RNN model for graph perturbation.
[0121] Proceeding to operation 1103, the proposed system is trained using NGP learning. The process includes generating a scene graph and utilizing an external knowledge base 1102, a dataset 1104 consisting of training images and scene graphs, and a symbolic theory 1105.
[0122] Operation 1106 indicates that the proposed system utilizes a pure neural AI model without any symbolic components in the perturbation.
[0123] At operation 1107, the AI model is compressed and quantized so that it can be deployed on a device.
[0124] At operation 1109, given an image or video frame 1108 in the real world or the meta-world, a deep neural network is optimized for the device hardware.
[0125] Operation 1110 indicates that a scene graph is being generated, where, at operation 1111, the proposed system generates a description based on facts according to the scene graph.
[0126] Finally, operation 1112 involves the proposed system narrating via Bixby or a controller, resulting in further actions.
[0127] The various actions, behaviors, blocks, operations, etc. in the flowchart can be performed in the presented order, in a different order, or simultaneously. Additionally, in some embodiments, some of the actions, behaviors, blocks, operations, etc. can be omitted, added, modified, skipped, etc. without departing from the scope of the present disclosure.
[0128] The proposed method and electronic device 300 have the following advantages: Privacy: The knowledge graph is not only personalized but also secure, and the parameters of the AI are fine-tuned so as not to disclose the user's identity.
[0129] Personalization: The proposed method and electronic device 300 are trained using a knowledge graph that simulates common sense. The knowledge graph is personalized, allowing the NSAI to be easily fine-tuned to meet the specific needs of the user.
[0130] Immediacy: The proposed method and electronic device 300 do not require storing the generated descriptive facts. This allows the NGP-based NSAI to be extended for versatile use.
[0131] Interpretable AI: The proposed method and electronic device 300 allow the AI to have the ability to grasp abstract concepts and explain behaviors in a human-like manner.
[0132] Accuracy: The present disclosure introduces a novel technology and electronic device 300 that improves the accuracy of traditional AI without any additional overhead.
[0133] Quickness: The method and electronic device 300 are designed to be converted into a deep neural network that is carefully optimized for fast execution on hardware.
[0134] Lightweight: The proposed method and electronic device 300 eliminate the need for the knowledge graph during interference, thus reducing the memory footprint.
[0135] Although the present disclosure has been shown and described with reference to various embodiments, those skilled in the art will understand that various changes in form and detail can be made therein without departing from the spirit and scope of the present disclosure as defined by the appended claims and their equivalents.
Claims
1. A method for neuro-symbolic learning of an artificial intelligence (AI) model, the method comprising: Receiving, by an electronic device, input data including a plurality of contents for the neuro-symbolic learning of the AI model; Determining, by the electronic device, a prediction probability for each of the plurality of contents in the output of the AI model for the input data; Determining, by the electronic device, a neural loss of the AI model by comparing the prediction probability for each of the plurality of contents with a predefined expected probability for each of the plurality of contents; Determining, by the electronic device, a symbolic loss of the AI model by comparing the prediction probability for each of the plurality of contents with a predetermined undesired probability for each of the plurality of contents; Determining, by the electronic device, weights of multiple layers of the AI model; And Updating, by the electronic device, the weights of the multiple layers of the AI model based on the neural loss and the symbolic loss.
2. The method according to claim 1, wherein The method for updating, by the electronic device, the weights of the multiple layers of the AI model based on the neural loss and the symbolic loss includes: Determining, by the electronic device, a training loss of the AI model as a measure of the neural loss and the symbolic loss; and Proportionally updating, by the electronic device, the weights of the multiple layers of the AI model according to the determined training loss.
3. The method according to claim 1, Among them, Determining, by the electronic device, a predetermined undesired probability for each of the plurality of contents in the input data, and wherein the method for determining a predetermined undesired probability for each of the plurality of contents includes: Selecting, by the electronic device, an external symbolic knowledge graph including a plurality of common sense facts; Constructing, by the electronic device, a negative knowledge graph including a plurality of facts violating common sense from the external symbolic knowledge graph; Identifying, by the electronic device, a set of the most violated facts from the plurality of facts violating common sense; and Generating, by the electronic device, a symbolic label in the form of a probability for the most violated facts, and wherein the most violated facts include the undesired probability.
4. The method according to claim 3, wherein The plurality of common sense facts includes at least one of a predefined rule set or a set of common sense facts associated with the real world.
5. The method according to claim 3, wherein, The most violated facts further include facts contrary to the plurality of common sense facts about the input data.
6. The method according to claim 1, further comprising: Determining, by the electronic device, the plurality of contents and the relationships between the plurality of contents; And Determining, by the electronic device, at least one scene graph based on the plurality of contents and the relationships between the plurality of contents, wherein the at least one scene graph includes a structural representation of the plurality of contents in the input data and the relationships between the plurality of contents.
7. The method according to claim 6, wherein Using neural AI to determine the plurality of contents, and using symbolic AI to determine the relationships between the plurality of contents, wherein the symbolic AI follows the common sense of the real world through an external symbolic knowledge graph.
8. An electronic device for neuro-symbolic learning of an artificial intelligence (AI) model, the electronic device comprising: A processor; A communicator; A neuro-symbolic AI controller; And A memory storing one or more programs including computer-executable instructions that, when run by the processor, cause the electronic device to perform the following operations: Receiving input data including a plurality of contents for the neuro-symbolic learning of the AI model, Determining a prediction probability for each of the plurality of contents in the output of the AI model for the input data, Determining a neural loss of the AI model by comparing the prediction probability for each of the plurality of contents with a predefined desired probability for each of the plurality of contents, Determining a symbolic loss of the AI model by comparing the prediction probability for each of the plurality of contents with a predefined undesired probability for each of the plurality of contents, Determining a training loss of the AI model as a measure of the neural loss and the symbolic loss, Determining weights of multiple layers of the AI model, and Updating the weights of the multiple layers of the AI model based on the neural loss and the symbolic loss.
9. The electronic device according to claim 8, wherein, The one or more programs further include instructions that, when run by the processor, cause the electronic device to perform the following operations: Determining a training loss of the AI model as a measure of the neural loss and the symbolic loss, and Proportionally updating the weights of the multiple layers of the AI model according to the determined training loss.
10. The electronic device according to claim 8, Among them, The one or more programs further include instructions that, when run by the processor, cause the electronic device to perform the following operations: Selecting an external symbolic knowledge graph including a plurality of common-sense facts, Constructing a negative knowledge graph including a plurality of counter common-sense facts from the external symbolic knowledge graph, Identifying a set of the most violated facts from the plurality of counter common-sense facts, and Generating a symbolic label for the most violated facts, Wherein the most violated facts are undesired, and Wherein the symbolic label is provided in the form of a probability.
11. The electronic device according to claim 10, wherein, The plurality of common-sense facts include at least one of a predefined rule set or a set of common-sense facts associated with the real world.
12. The electronic device according to claim 10, wherein, The most violated facts include facts that are contrary to the plurality of common-sense facts regarding the input data.
13. The electronic device according to claim 8, Among them, The one or more programs further include instructions that, when run by the processor, cause the electronic device to perform the following operations: Determining the plurality of contents and the relationships between the plurality of contents, and Determining at least one scene graph based on the determined plurality of contents and the relationships between the plurality of contents, and Wherein the at least one scene graph includes a structural representation of the plurality of contents in the input data and the relationships between the plurality of contents.
14. The electronic device according to claim 13, wherein, Use neural AI to determine the multiple contents, and use symbolic AI to determine the relationships between the multiple contents, wherein the symbolic AI follows real-world common sense through an external symbolic knowledge graph.
15. The electronic device according to claim 8, Among them, determine the training loss by applying various functions to both the neural loss and the symbolic loss, and wherein the various functions include addition, subtraction, multiplication, and division.