Method and device for realizing logical reasoning based on computing power of intelligent computing center
By running a logical inference model on the intelligent computing center, object and scene recognition and logical inference are performed on the image, the problem of lack of efficient logical inference in image evaluation in the prior art is solved, and high-precision and fast-responsive image evaluation is achieved.
Patent Information
- Application Number
- CN202510223937.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art lacks efficient logical reasoning methods in image evaluation, and it is difficult to meet the requirements of real-time and high-precision, especially when dealing with complex and multi-dimensional logical reasoning tasks.
The logical reasoning method is implemented based on the computing power based on the intelligent computing center. By obtaining the image to be evaluated, the logical reasoning model running in the intelligent computing center infers the image, identifying objects and scenes, extracting target features, and inferring the logical relationship between objects, scenes, and objects and scenes based on these features.
It improves the accuracy and efficiency of image logic reasoning, can respond quickly to image evaluation, and solves the problems of inconsistent evaluation standards and limited intelligence in the prior art.
Smart Images

Figure CN120070919A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent computing centers and computing power infrastructure of intelligent computing centers, and particularly relates to a method and device for realizing logical reasoning based on the computing power of an intelligent computing center. Background Art
[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.
[0003] An "intelligent computing center" refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.), and mainly provides the required computing power, data, and algorithms for artificial intelligence applications (such as artificial intelligence deep learning model development, model training, and model inference scenarios). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0004] The "intelligent computing center" includes, but is not limited to, the "intelligent computing center".
[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services, and algorithm services required for artificial intelligence applications.
[0006] "Computing power" is the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of a target result by processing information data, and a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, which mainly provides services to society through computing power infrastructure.
[0007] With the continuous development of generative artificial intelligence technology, the application of image generation and processing technology in the fields of creative design, virtual reality, and cultural inheritance has gradually increased. However, there is still a lack of a unified standard and an efficient implementation method for the evaluation method of generated pictures. Traditional image evaluation methods do not utilize intelligent computing centers and computing power, and are limited by computing power and algorithm capabilities. When dealing with complex and multi-dimensional logical reasoning tasks, it is difficult to meet the requirements of real-time performance and high precision; and the existing technologies mainly focus on the analysis of the explicit features (such as color, shape, position, etc.) of image content, but lack an effective evaluation method for the logicality and semantic rationality of image content. For example, in the scenario of "a little monk holding a broom", the generated picture needs to reflect a reasonable background and context (such as a temple scene), rather than an environment that violates semantic logic (such as a city road). The existing technologies are difficult to comprehensively consider the semantic association between objects and scenes in the image, resulting in inconsistent evaluation criteria and limited intelligence, and it is difficult to meet the needs of high-level semantic reasoning. Summary of the Invention
[0008] The present invention provides a method and device for realizing logical reasoning based on the computing power of an intelligent computing center to solve the problem of the lack of an efficient logical reasoning method in existing image evaluation.
[0009] To solve the above technical problems, the present invention is implemented as follows:
[0010] In a first aspect, the present invention provides a method for realizing logical reasoning based on the computing power of an intelligent computing center, including:
[0011] Step S1: Obtain an image to be evaluated;
[0012] Step S2: Use a logical reasoning model running on the intelligent computing center to reason about the image to be evaluated, and obtain the objects, scenes, and logical information between the objects and scenes in the image to be evaluated; wherein, the logical reasoning model reasoning about the image to be evaluated includes: identifying the objects and scenes in the image to be evaluated; extracting the target features of the objects and scenes in the image to be evaluated; reasoning about the logical relationships between the objects, scenes, and between the objects and scenes according to the target features, and obtaining the objects, scenes, and logical information between the objects and scenes in the image to be evaluated.
[0013] Optionally, step S1 includes:
[0014] Step S11: Obtain an original image to be evaluated;
[0015] Step S12: Preprocess the original image to be evaluated to obtain the image to be evaluated, and the preprocessing includes at least one of the following: adjusting the size, normalizing, enhancing data features, and denoising.
[0016] Optionally, the logical information includes at least one of the following: visual element logic, image composition logic, image syntax logic, color logic, space logic, narrative structure, cultural background, and social background logic.
[0017] Optionally, before step S2, it further includes:
[0018] Step S0: Train the logical reasoning model to be trained;
[0019] Step S0 includes:
[0020] Step S01: Obtain training images and the real logical information between the objects, scenes, and between the objects and scenes in the training images;
[0021] Step S02: Use the logic inference model to be trained to infer the training image, and obtain the predicted objects, scenes, and logical information between the objects and scenes in the training image;
[0022] Step S03: Optimize the logic inference model to be trained according to the predicted objects, scenes, and logical information between the objects and scenes in the training image and the true logical information between the objects, scenes, and objects and scenes in the training image, so as to obtain the trained logic inference model.
[0023] Optionally, after step S2, the following is further included:
[0024] Step S3: Visually output the objects, scenes, and logical information between the objects and scenes in the image to be evaluated in the form of image annotation or text.
[0025] Optionally, after step S2, the following is further included:
[0026] Step S4: Perform similarity matching between the objects, scenes, and logical information between the objects and scenes in the image to be evaluated and the logical information stored in advance when generating the image to be evaluated, so as to obtain the inference score of the logical elements.
[0027] In a second aspect, the present invention provides a logic inference device based on the computing power of an intelligent computing center, including:
[0028] An acquisition module, configured to acquire an image to be evaluated;
[0029] A processing module, configured to use the logic inference model running on the intelligent computing center to infer the image to be evaluated, and obtain the objects, scenes, and logical information between the objects and scenes in the image to be evaluated; wherein, the logic inference model infers the image to be evaluated including: identifying the objects and scenes in the image to be evaluated; extracting the target features of the objects and scenes in the image to be evaluated; and inferring the logical relationship between the objects, scenes, and objects and scenes according to the target features, so as to obtain the objects, scenes, and logical information between the objects and scenes in the image to be evaluated.
[0030] Optionally, the acquisition module includes:
[0031] A first acquisition sub-module, configured to acquire the original image to be evaluated;
[0032] A preprocessing sub-module, configured to preprocess the original image to be evaluated to obtain the image to be evaluated, and the preprocessing includes at least one of the following: adjusting the size, normalizing, enhancing data features, and denoising.
[0033] Optionally, the logical information includes at least one of the following: visual element logic, image composition logic, image grammar logic, color logic, spatial logic, narrative structure, cultural background, and social background logic.
[0034] Optionally, it further includes:
[0035] A model training module for training a logical reasoning model to be trained;
[0036] The model training module includes:
[0037] A first acquisition sub-module for acquiring training images and the objects, scenes, and the true logical information between the objects and scenes in the training images;
[0038] An inference sub-module for inferring the training images using the logical reasoning model to be trained, and obtaining the predicted logical information between the objects, scenes, and the objects and scenes in the training images;
[0039] An optimization sub-module for optimizing the logical reasoning model to be trained according to the predicted logical information between the objects, scenes, and the objects and scenes in the training images and the true logical information between the objects, scenes, and the objects and scenes in the training images, so as to obtain a trained logical reasoning model.
[0040] Optionally, it further includes:
[0041] A visualization output module for visually outputting the logical information between the objects, scenes, and the objects and scenes in the image to be evaluated in the form of image annotation or text.
[0042] Optionally, it further includes:
[0043] A scoring module for performing similarity matching on the logical information between the objects, scenes, and the objects and scenes in the image to be evaluated and the logical information stored in advance when generating the image to be evaluated, so as to obtain an inference score for the logical elements.
[0044] In a third aspect, the present invention provides an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements the steps in the logical reasoning method based on the computing power of the intelligent computing center as described in any item of the first aspect.
[0045] In a fourth aspect, the present invention provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements the steps in the logical reasoning method based on the computing power of the intelligent computing center as described in any item of the first aspect.
[0046] In a fifth aspect, the present invention provides a computer program product including computer instructions which, when executed by a processor, implement the steps in the logical reasoning method based on the computing power of an intelligent computing center as described in any one of the first aspects.
[0047] In the present invention, an image to be evaluated is obtained; a logical reasoning model running in the intelligent computing center is used to perform reasoning on the image to be evaluated, and object, scene, and logical information between the object and the scene in the image to be evaluated are obtained; wherein, the logical reasoning model performing reasoning on the image to be evaluated includes: identifying the object and the scene in the image to be evaluated; extracting target features of the object and the scene in the image to be evaluated; and performing reasoning on the object, the scene, and the logical relationship between the object and the scene according to the target features, so as to obtain the object, the scene, and the logical information between the object and the scene in the image to be evaluated. In this way, through the powerful computing power provided by the intelligent computing center, the logical reasoning model is used to intelligently identify and reason about the logic and semantic rationality of the image content, improving the accuracy and efficiency of image logical reasoning, thereby enabling a quick response to image evaluation and solving the problem of the lack of an efficient logical reasoning method in existing image evaluation. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of illustrating the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0049] Figure 1 is a flowchart of a logical reasoning method based on the computing power of an intelligent computing center provided by the present invention;
[0050] Figure 2 is a schematic structural diagram of a logical reasoning device based on the computing power of an intelligent computing center provided by the present invention;
[0051] Figure 3 is a schematic structural diagram of an electronic device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] The technical solutions in the present invention will be clearly and completely described below with reference to the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.
[0053] The "computing power" described in the present invention refers to the ability of computer devices or computing / data centers to process information, which is the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement. It is the computing ability to achieve the output of the target result through the processing of information data. It is a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.
[0054] The "computational power" (Computational Power, CP) described in the present invention is an ability of the data center server to process data and achieve result output, and it is a comprehensive indicator to measure the computing ability of the data center, including general computing ability, supercomputing ability, and intelligent computing ability. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1 EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing ability. It is estimated that 1 EFLOPS is approximately the computing power output of 5 Tianhe 2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP_general + CP_intelligent + CP_super.
[0055] The "network power" (Network Power, NP) described in the present invention is the manifestation of the data transmission ability of the computing power facility, and it is a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., which involves network transmission inside and between data centers and is a comprehensive indicator to measure the network transmission scheduling ability.
[0056] The "storage power" (Storage Power, SP) described in the present invention is the comprehensive ability of the data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon. It is a comprehensive indicator to measure the data storage ability of the data center, including external storage devices such as storage arrays and server internal storage devices. The commonly used measurement unit for storage capacity is exabyte (EB, 1 EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read and write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.
[0057] The "computing power infrastructure" described in the present invention refers to: a new type of information infrastructure integrating information computing power, network carrying capacity, and data storage capacity, which can realize the centralized computing, storage, transmission, and application of information, presenting characteristics such as multi-element ubiquitous, intelligent and agile, secure and reliable, and green and low-carbon.
[0058] The "new information infrastructure" described in the present invention refers to: mainly including network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, and satellite Internet, computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, and supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing. With the emergence and popularization of new general technologies, the form of the new information infrastructure will be more diverse.
[0059] The "computing power" described in the present invention includes: general computing power, intelligent computing power, and super computing power.
[0060] The "general computing power" described in the present invention refers to: the computing power provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0061] The "intelligent computing power" described in the present invention refers to: for various artificial intelligence innovation applications, a computing platform deployed on a large scale based on dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit), such as natural language processing, machine vision, and so on.
[0062] The "super computing power" described in the present invention refers to: mainly the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system, mainly for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.
[0063] The "intelligent computing center" described in the present invention refers to: a facility that mainly provides the required computing power, data, and algorithms for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0064] The "intelligent computing center" described in the present invention includes but is not limited to the "intelligent computing center".
[0065] The "intelligent computing center" described in the present invention, namely the artificial intelligence computing center, is a type of computing power infrastructure that provides computing power services, data services, and algorithm services required for artificial intelligence applications based on artificial intelligence theory and adopting an artificial intelligence computing architecture.
[0066] The "computing power center" described in the present invention refers to: a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, with computing power, carrying capacity, and storage capacity, including general data centers, intelligent computing centers, supercomputing centers, etc.
[0067] The "supercomputing center" described in the present invention refers to: that is, a supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters, and can provide functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.
[0068] The "computing power resources" described in the present invention refers to: technologies and facilities with information computing, transmission, storage, and application capabilities required for the development of the digital society, including but not limited to computing resources such as CPUs and GPUs, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.
[0069] The "logic" described in the present invention refers to: generating the semantic rationality and consistency of objects, scenes, and their relationships in the picture, that is, whether the image content conforms to human daily cognition and common sense rules. Logic not only includes the semantic correctness of the image content, but also covers the internal connections between objects, scenes, and behaviors, as well as the contextual coordination of the overall scene.
[0070] The "logical reasoning" described in the present invention refers to: in the evaluation process of generating pictures, a process of automatically analyzing and verifying the logical consistency and rationality of the image content based on semantic information, common sense knowledge, and context relationships. Through logical reasoning, it can be judged whether there is an association that conforms to common sense or a specific context between objects, behaviors, and scenes in the image, so as to ensure the semantic and logical correctness of the generated pictures.
[0071] Please refer to Figure 1 , the present invention provides a method for realizing logical reasoning based on the computing power of an intelligent computing center, including:
[0072] Step S1: Obtain the image to be evaluated;
[0073] In the present invention, optionally, the step S1 includes:
[0074] Step S11: Obtain the original image to be evaluated;
[0075] Step S12: Preprocess the original image to be evaluated to obtain the image to be evaluated, and the preprocessing includes at least one of the following: resizing, normalization processing, data feature enhancement, and denoising.
[0076] In the present invention, the original image to be evaluated is preprocessed to provide clearer and more standardized image data for subsequent analysis and processing, improve the quality and consistency of the image data, and thus enhance the effect of subsequent processing and analysis.
[0077] In the present invention, the image to be evaluated can be an image directly uploaded or an image generated by artificial intelligence. The image to be evaluated includes different objects or elements, as well as corresponding scenes. There are interactions and connections between different objects or elements in the image, and between objects or elements and the scene. Specifically, logic refers to the semantic rationality and consistency of the objects, scenes, and their relationships in the generated picture, that is, whether the image content conforms to human daily cognition and common sense rules. Logic not only includes the semantic correctness of the image content, but also covers the internal connections between objects, scenes, and behaviors, as well as the situational coordination of the overall scene.
[0078] Step S2: Use the logical reasoning model running in the intelligent computing center to reason about the image to be evaluated, and obtain the objects, scenes, and logical information between the objects and the scene in the image to be evaluated; wherein, the logical reasoning model reasoning about the image to be evaluated includes: identifying the objects and scenes in the image to be evaluated; extracting the target features of the objects and scenes in the image to be evaluated; reasoning about the logical relationships between the objects, scenes, and between the objects and the scene according to the target features, and obtaining the objects, scenes, and logical information between the objects and the scene in the image to be evaluated.
[0079] In the present invention, optionally, before the step S2, it further includes:
[0080] Step S0: Perform model training on the logical reasoning model to be trained;
[0081] The step S0 includes:
[0082] Step S01: Obtain training images and the real logical information between the objects, scenes, and between the objects and the scene in the training images;
[0083] Step S02: Use the logical reasoning model to be trained to reason about the training images, and obtain the predicted logical information between the objects, scenes, and between the objects and the scene in the training images;
[0084] Step S03: Optimize the logical reasoning model to be trained according to the predicted logical information between the objects, scenes, and between the objects and the scene in the training images and the real logical information between the objects, scenes, and between the objects and the scene in the training images, and obtain the trained logical reasoning model.
[0085] In the present invention, the logical reasoning model is trained using training images, and the parameters of the model are adjusted to minimize the loss function. Among them, common optimization algorithms can be used but are not limited to gradient descent or Adaptive Moment Estimation (Adam), etc., and the performance of the model can be monitored, and the model can be updated regularly to cope with changes in data distribution or the introduction of new data. Among them, the logical reasoning model can, but is not limited to, use a pre-trained visual model to extract the target features of the objects in the image to be evaluated, and combine object detection and logical reasoning to identify the specific logic between the objects and the scene in the picture.
[0086] In the present invention, the logical reasoning model reasons about the logical consistency between objects and scenes at the semantic level. For example, in the scene of "a little monk holding a broom", the system can identify objects such as "little monk" and "broom", and combine background information for logical reasoning to conclude that the little monk should be in the temple yard, and judge whether the image conforms to semantic rationality. The system first analyzes the objects and scenes in the image in layers, and then judges whether the content of the picture conforms to real semantics and common sense logic through semantic matching and logical rule reasoning.
[0087] In the present invention, optionally, the logical information includes at least one of the following: visual element logic, image composition logic, image grammar logic, color logic, spatial logic, narrative structure, cultural background, and social background logic.
[0088] In the present invention, a logical reasoning model running in the intelligent computing center is adopted. Through the powerful computing power provided by the intelligent computing center, it supports efficient reasoning in complex scenarios and can flexibly adapt to different application scenarios. And a logical reasoning model is adopted to meet different image evaluation requirements.
[0089] Specifically, the visual element logic includes basic elements such as shapes, colors, lines, textures, etc. in the image, and how they are combined and interact with each other to form an overall visual effect; the image composition logic is the layout and structure of the image, including the arrangement, symmetry, balance, and focus of elements, etc., which affect the viewer's visual flow and attention; the image grammar logic is the visual language rules of the image, including how to convey information through visual elements, similar to the grammar rules in language; the color logic is how the selection and combination of colors affect the emotional expression and visual attraction of the image; the spatial logic is how to represent depth and spatial relationships in the image, affecting the viewer's understanding and feeling of the image; the narrative structure is how the image tells a story or conveys information, and the temporal and action relationships in the image; the cultural background and social background logic are that the logic of the image is also affected by cultural and social backgrounds, and different cultures may interpret images differently. The logical information constitutes the logic of the image and can express the information and emotions conveyed by the image.
[0090] In the present invention, a to-be-evaluated image is obtained; a logical reasoning model running in the intelligent computing center is used to perform reasoning on the to-be-evaluated image to obtain the objects, scenes, and logical information between the objects and scenes in the to-be-evaluated image; wherein, the logical reasoning model performing reasoning on the to-be-evaluated image includes: identifying the objects and scenes in the to-be-evaluated image; extracting the target features of the objects and scenes in the to-be-evaluated image; and performing reasoning on the objects, scenes, and logical relationships between the objects and scenes according to the target features to obtain the objects, scenes, and logical information between the objects and scenes in the to-be-evaluated image. In this way, through the powerful computing power provided by the intelligent computing center, the logical reasoning model is used to intelligently identify and reason about the logic and semantic rationality of the image content, improving the accuracy and efficiency of image logical reasoning, so as to perform a quick response to image evaluation and solve the problem of the lack of an efficient logical reasoning method in the existing image evaluation.
[0091] In the present invention, optionally, after step S2, the following is further included:
[0092] Step S3: Visualize and output the objects, scenes, and logical information between the objects and scenes in the to-be-evaluated image in the form of image annotation or text.
[0093] In the present invention, outputting the reasoning result, that is, the objects, scenes, and logical information between the objects and scenes in the to-be-evaluated image in a visual manner, includes an image annotating the logical information of the objects, or describing the objects, scenes, and logical information between the objects and scenes in the image in the form of text or charts, improving the transparency and usability of the model, and being able to more clearly display the evaluation result of the to-be-evaluated image, facilitating further evaluation by subsequent manual or system, so as to promote better business results.
[0094] In the present invention, optionally, after step S2, the following is further included:
[0095] Step S4: Perform similarity matching on the objects, scenes, and logical information between the objects and scenes in the to-be-evaluated image and the logical information stored in advance when generating the to-be-evaluated image to obtain the reasoning score of the element logic.
[0096] In the present invention, the logical reasoning model can, but is not limited to, use a reward model to perform similarity matching on the objects, scenes, and logical information between the objects and scenes in the to-be-evaluated image and the logical information stored in advance when generating the to-be-evaluated image to obtain the reasoning score of the element logic;
[0097] In some embodiments, taking the example that the image to be evaluated is from an Artificial Intelligence Generated Content (AIGC) generation model, the reward model optimizes the framework of the output of the generation model through a reward signal. That is, in the evaluation task of the image to be evaluated generated by AIGC, the goal of the reward model is to automatically measure the quality of the image to be evaluated and provide meaningful feedback to guide the generation model (such as a diffusion model) to generate higher-quality pictures. Specifically, the quality evaluation here is mainly based on the reasoning ability of the logical information between objects or elements in the picture and the scene. In other words, the task of the reward model is to analyze whether the logical information between the objects and the scene in the generated picture meets the expectations and score or reward based on this.
[0098] Specifically, by calculating the similarity between the object, scene, and the logical information A between the object and the scene in the image to be evaluated and the logical information B stored in advance when generating the image to be evaluated: Rattr = Sim(A, B); where Sim(·) can use cosine similarity, Euclidean distance, or other feature comparison methods;
[0099] Then, use the semantic similarity score between the text and the picture to obtain the multi-modal consistency reward: Rtext-img = Sim(CLIPimg(image), CLIPtext(input)); where CLIPimg(image) is the first feature vector obtained by inputting the image image into the image encoder of the multi-modal pre-training model (Contrastive Language-Image Pre-training, CLIP), representing the semantic information of the picture; CLIPtext(input) is the second feature vector obtained by inputting the text input into the text encoder of CLIP, representing the semantic information of the text.
[0100] Traditional image quality evaluation metrics can be introduced, such as Peak Signal-to-Noise Ratio (PSNR), Structural Similarity (SSIM), or a dedicated network can be trained to predict the subjective quality score of the picture, so as to integrate various evaluation metrics to calculate the final reward score. In the subsequent reinforcement learning stage, the loss function for guiding the improvement of the generation model through the score obtained from the reward model is: LossRL = -Epolicy[R(image)], where Epolicy[R(image)] is the reward score obtained for the image generated by the model under the current policy (policy).
[0101] In the present invention, by designing a reward function, the quality of picture pieces is comprehensively evaluated from aspects such as the matching degree between logical information, multimodal consistency, and visual quality dimension, and a reward model is trained through supervised learning. The reward model is combined with a generation model, and reinforcement learning is used to optimize the generation process to improve the quality, attribute consistency, and diversity of picture generation.
[0102] Please refer to Figure 2 , an embodiment of the present invention provides a logical reasoning device based on the computing power of an intelligent computing center, including:
[0103] An acquisition module 21, configured to acquire an image to be evaluated;
[0104] A processing module 22, configured to perform reasoning on the image to be evaluated by using a logical reasoning model running on the intelligent computing center to obtain objects, scenes, and logical information between the objects and the scenes in the image to be evaluated; wherein, the logical reasoning model performing reasoning on the image to be evaluated includes: identifying the objects and scenes in the image to be evaluated; extracting target features of the objects and scenes in the image to be evaluated; reasoning on the logical relationship between the objects, the scenes, and between the objects and the scenes according to the target features to obtain the objects, the scenes, and the logical information between the objects and the scenes in the image to be evaluated.
[0105] In the present invention, optionally, the acquisition module includes:
[0106] A first acquisition sub-module, configured to acquire an original image to be evaluated;
[0107] A preprocessing sub-module, configured to preprocess the original image to be evaluated to obtain the image to be evaluated, and the preprocessing includes at least one of the following: adjusting the size, normalization processing, data feature enhancement, and denoising.
[0108] In the present invention, optionally, the logical information includes at least one of the following: visual element logic, image composition logic, image grammar logic, color logic, space logic, narrative structure, cultural background, and social background logic.
[0109] In the present invention, optionally, it further includes:
[0110] A model training module, configured to perform model training on a logical reasoning model to be trained;
[0111] The model training module includes:
[0112] A first acquisition sub-module, configured to acquire training images and the true logical information between the objects, the scenes, and between the objects and the scenes in the training images;
[0113] An inference sub-module, configured to perform inference on the training image by using a logic inference model to be trained, and obtain the predicted objects, scenes, and the logical information between the objects and the scenes in the training image;
[0114] An optimization sub-module, configured to optimize the logic inference model to be trained according to the predicted objects, scenes, and the logical information between the objects and the scenes in the training image and the true logical information between the objects, scenes, and the objects and the scenes in the training image, so as to obtain a trained logic inference model.
[0115] In the present invention, optionally, it further includes:
[0116] A visualization output module, configured to visually output the objects, scenes, and the logical information between the objects and the scenes in the image to be evaluated in the form of image annotation or text.
[0117] In the present invention, optionally, it further includes:
[0118] A scoring module, configured to perform similarity matching on the objects, scenes, and the logical information between the objects and the scenes in the image to be evaluated and the logical information stored in advance when generating the image to be evaluated, so as to obtain an inference score of the logical elements.
[0119] The logic inference based on the computing power of the intelligent computing center provided by the embodiments of the present invention can implement Figure 1 each process implemented by the method embodiments, and achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0120] The embodiments of the present invention provide an electronic device 30. Refer to Figure 3 as shown Figure 3 which is a schematic block diagram of the electronic device 30 according to the embodiments of the present invention, including a processor 31, a memory 32, and a program or instruction stored in the memory 32 and executable on the processor 31. When the program or instruction is executed by the processor, it implements the steps in any one of the logic inference methods based on the computing power of the intelligent computing center of the present invention.
[0121] The embodiments of the present invention provide a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the embodiments of the logic inference method based on the computing power of the intelligent computing center as described above, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0122] The embodiments of the present invention further provide a computer program product, including computer instructions, which when executed by a processor, implement the above Figure 1 each process of the method embodiments shown, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0123] A computer-readable medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. As defined in the present invention, a computer-readable medium does not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0124] It should be noted that in the technical solution of the present invention, in aspects such as the collection, gathering, updating, analysis, processing, use, transmission, and storage of the user's personal information, it complies with the provisions of relevant laws and regulations, is used for legal purposes, and does not violate public order and good customs. Necessary measures are taken for the user's personal information to prevent illegal access to the user's personal information data and to maintain the security of the user's personal information and network security.
[0125] It should be noted that in the present invention, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising that element.
[0126] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0127] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a service classification device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0128] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A method for implementing logical reasoning based on the computing power of an intelligent computing center, characterized in that: include: Step S1: Obtain an image to be evaluated; Step S2: Use the logical reasoning model running in the intelligent computing center to reason about the image to be evaluated, and obtain the objects, scenes, and logical information between the objects and scenes in the image to be evaluated; wherein the logical reasoning model reasoning about the image to be evaluated includes: identifying the objects and scenes in the image to be evaluated; extracting the target features of the objects and scenes in the image to be evaluated; reasoning about the objects, scenes, and the logical relationships between the objects and scenes based on the target features, and obtain the objects, scenes, and logical information between the objects and scenes in the image to be evaluated.
2. The method for realizing logical reasoning based on the computing power of an intelligent computing center according to claim 1, characterized in that: The step S1 comprises: Step S11: obtaining the original image to be evaluated; Step S12: preprocessing the original image to be evaluated to obtain the image to be evaluated, wherein the preprocessing includes at least one of the following: resizing, normalization, data feature enhancement, and denoising.
3. The method for realizing logical reasoning based on the computing power of an intelligent computing center according to claim 1, characterized in that: The logic information includes at least one of the following: visual element logic, image composition logic, image grammar logic, color logic, space logic, narrative structure, cultural background and social background logic.
4. The method for realizing logical reasoning based on the computing power of an intelligent computing center according to claim 1, characterized in that: Before step S2, the method further includes: Step S0: Perform model training on the logical reasoning model to be trained; The step S0 comprises: Step S01: obtaining a training image and objects, scenes in the training image, and real logical information between the objects and scenes; Step S02: using the logic reasoning model to be trained to reason on the training image to obtain predicted objects, scenes, and logical information between the objects and scenes in the training image; Step S03: Optimize the logical reasoning model to be trained based on the predicted logical information between the objects, scenes and the objects in the training image and the actual logical information between the objects, scenes and the objects in the training image to obtain a trained logical reasoning model.
5. The method for realizing logical reasoning based on the computing power of an intelligent computing center according to claim 1, characterized in that: The step S2 further includes: Step S3: Visually output the objects, scenes, and logical information between the objects and scenes in the image to be evaluated in the form of image annotations or text.
6. The method for realizing logical reasoning based on the computing power of an intelligent computing center according to claim 1, characterized in that: The step S2 further includes: Step S4: performing similarity matching on the objects, scenes, and logical information between the objects and scenes in the image to be evaluated with the pre-stored logical information stored when the image to be evaluated is generated, to obtain an inference score of the logic of the elements in the image.
7. A logical reasoning device based on the computing power of an intelligent computing center, characterized in that: include: An acquisition module, used for acquiring an image to be evaluated; A processing module is used to use the logical reasoning model running in the intelligent computing center to reason about the image to be evaluated, so as to obtain the objects, scenes and logical information between the objects and scenes in the image to be evaluated; wherein the logical reasoning model reasoning about the image to be evaluated includes: identifying the objects and scenes in the image to be evaluated; extracting the target features of the objects and scenes in the image to be evaluated; and reasoning about the objects, scenes and the logical relationships between the objects and scenes according to the target features to obtain the objects, scenes and logical information between the objects and scenes in the image to be evaluated.
8. An electronic device, characterized in that: It includes a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps in the method for realizing logical reasoning based on the computing power of an intelligent computing center as described in any one of claims 1 to 6.
9. A readable storage medium, characterized in that: The readable storage medium stores programs or instructions, and when the programs or instructions are executed by the processor, the steps in the method for realizing logical reasoning based on the computing power of an intelligent computing center as described in any one of claims 1 to 6 are implemented.
10. A computer program product, characterized in that It includes computer instructions, which, when executed by a processor, implement the steps in the method for realizing logical reasoning based on the computing power of an intelligent computing center as described in any one of claims 1 to 6.