Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

172 results about "Object description" patented technology

Audio and video player control method based on voice instruction

The invention relates to the technical field of audio and video control, and discloses an audio and video player control method based on a voice instruction. The method comprises the steps that an original voice instruction stream of a user is collected, the instruction stream comprises a time domain audio signal sequence, an environment noise spectrum and user pronunciation characteristic parameters, and voice information can be comprehensively captured; multi-modal instruction analysis processing is carried out on the original voice instruction stream, a structured control instruction set containing acoustic control intention identification, semantic operation object description and context correlation parameters is generated, and the analysis precision is improved; then executing player state adaptation based on the set, generating a dynamic control response sequence containing an equipment state adjustment command, a media content positioning parameter and an interface interaction logic identifier, driving a player to execute a multi-dimensional control operation and generating real-time play control effect feedback data; and finally, multi-modal analysis parameters are optimized according to feedback data, a self-adaptive instruction analysis strategy is generated, and the control experience of a user on the audio and video player is optimized.
Owner:ONWAY TECH LTD

Intelligent SQL (Structured Query Language) generation system based on multistage intention recognition and generation method thereof

The invention discloses an intelligent SQL (Structured Query Language) generation method and system based on multistage intention recognition. The method comprises the following steps: receiving a natural language query of a user; vector matching: carrying out vector matching based on a knowledge base to obtain Top-K candidate query objects, the knowledge base containing multilayer structure description information; a large language model intention understanding step: determining a final query object in combination with the candidate query object description; loading a branch workflow, analyzing a field and a table structure, and dynamically configuring a database table and field information; sQL generation: generating SQL statements conforming to grammatical rules, including automatic field addition, time condition conversion and field priority matching operation; and executing the SQL statement and returning a result, and if the result fails, recording a log to prompt correction. By combining a natural language processing technology with a database knowledge base, an efficient, accurate and user-friendly database query solution is provided.
Owner:SICHUAN ZHONGLI JIAHUA INFORMATION TECH CO LTD

Image object detection method, system and apparatus, and storage medium

Embodiments of the present description provide an image object detection method. The method comprises: on the basis of an image to be retrieved, an object description text, and an object retrieval condition, determining, by means of an object detection model, a target position of an object to be retrieved in said image, wherein the object description text is used for describing said object, and the object retrieval condition comprises at least one of a mask image, a pose, and a texture corresponding to said object.
Owner:ZHEJIANG DAHUA TECH CO LTD

Mandatory access control method and device based on process function context

The invention discloses a mandatory access control method and device based on a process function context, and the method comprises the steps: collecting a security context associated with a system call initiated by a target process, so as to generate a standardized object description; mapping the object description into a target function classification identifier, so as to obtain a process function context view of the target process according to the target function classification identifier; constructing a target decision key for access decision based on the current policy era, the qualifier, the function classification identifier, the view identifier of the process function context view and the isolation domain abstract; and querying the multi-level cache according to the target decision key to determine a matched target access decision. Therefore, context-sensitive judgment and cross-component consistency taking the functional context as the center are realized.
Owner:BEIJING METRO INFORMATION DEV CO LTD

Method of generating virtual avatar based on large model, agent, electronic device and storage medium

A method of generating a virtual avatar based on a large model, an agent, an electronic device and a storage medium, which relate to a field of artificial intelligence technology, and to fields of computer vision technology, deep learning technology, large model technology, etc., and may be applied to scenarios such as AIGC, digital character, intelligent e-commerce, etc. The method includes: processing a target image including a target object by using a large model to obtain object description information, the target object having texture information; processing the target image and a to-be-processed image representing an object morphology of a three-dimensional object by using a texture-generative large model to obtain a target three-dimensional object with target texture information, the three-dimensional object being determined based on the object description information, the target texture information being matched with the texture information; and generating the virtual avatar based on the target three-dimensional object.
Owner:BEIJING BAIDU NETCOM SCI & TECH CO LTD

Service processing

Object description information that is transmitted by a biometric recognition apparatus is received, the object description information includes a biometric feature of a target object and location information of the target object. Identity information of the target object is obtained according to the biometric feature. A service information set in association with the identity information is obtained. From the service information set, one or more pieces of candidate service information are selected. The one or more pieces of candidate service information are transmitted to a terminal device associated with the identity information. At least a first piece of target service information returned by the terminal device is received. At least a first service corresponding to the first piece of target service information is processed. Apparatus and non-transitory computer-readable storage medium counterpart embodiments are also contemplated.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Iterative automatic labeling of media data for artificial intelligence applications

Disclosed are apparatuses, systems, and techniques for automated iterative content detection and annotation of objects in media items. The techniques include performing a plurality of iterations to identify objects represented in a media item and referenced in a plurality of object descriptions of a prompt. An individual iteration includes identifying, using a content detection model, a subset of the objects represented in the media item and referenced in the plurality of object descriptions, or no objects represented in the media item and referenced in the plurality of object descriptions. Using the content detection model includes applying the content detection model to the media item and to the prompt or to an iteration prompt obtained from the prompt by eliminating descriptions of the subsets of the objects identified during previous iterations. The techniques further include generating, using the identified objects, a characterization of the media item.
Owner:NVIDIA CORP

Data processing method and device based on artificial intelligence, electronic equipment, computer readable storage medium and computer program product

The invention provides a data processing method and device based on artificial intelligence, electronic equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: acquiring an object grid model and a first texture parameter corresponding to the object grid model, and performing deformation processing on the object grid model based on a first weight parameter to obtain a first deformed object grid model; performing rendering processing based on the first deformation object grid model and the first texture parameter to obtain a first rendered object image, and performing image generation processing based on the object description text to obtain a generated object image conforming to the object description text; determining image semantic loss based on the first rendered object image and the generated object image, and updating the first weight parameter based on the image semantic loss to obtain a second weight parameter; and performing deformation processing on the object grid model based on the second weight parameter to obtain a second deformation object grid model. According to the invention, the modeling accuracy of the second deformation object grid model is improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD +1

Upside down reinforcement learning for text-to-image generation

A method, apparatus, non-transitory computer readable medium, and system for image processing include obtaining an input prompt including an image quality level and a description of an object, generating an image embedding based on the input prompt, where the image embedding represents the object and the image quality level in a vector space, and generating a synthetic image based on the image embedding, where the synthetic image depicts the object and has the image quality level.
Owner:ADOBE INC

Text-guided zero sample target counting method, program product and electronic equipment

The invention belongs to the technical field of image processing, and provides a text-guided zero sample target counting method, a program product and electronic equipment, and the counting method comprises the steps: inputting a query image and a target description text into a target counting model to obtain a density estimation graph of a target; the target counting model comprises a visual feature extraction module; a text feature extraction module; the multi-level polarity dual attention module comprises a plurality of cascaded feature fusion modules; each feature fusion module performs visual polarity cross attention processing on the visual-to-text fusion features or the visual features to obtain visual-to-text fusion features of the hierarchy, and performs text polarity cross attention processing on the text-to-visual fusion features or the text features to obtain text-to-visual fusion features of the hierarchy; the decoder is used for decoding the visual-to-text fusion features to obtain a density estimation graph of the target; according to the method, the information of the negative polarity part is reserved, and the accuracy of the density estimation graph of the target is improved.
Owner:CHONGQING UNIV

Live broadcast processing method and device, equipment and storage medium

The invention provides a live broadcast processing method and device, equipment and a storage medium, and the method comprises the steps: responding to a live broadcast plan generation operation, displaying to-be-recommended object information and live broadcast verbal skill information corresponding to a target live broadcast account number on a live broadcast plan page, the to-be-recommended object information comprising object attribute information of a to-be-recommended object, and the to-be-recommended object attribute information corresponding to the target live broadcast account number; the to-be-recommended object is determined according to historical interaction data corresponding to the target live broadcast account and object attribute information of the candidate to-be-recommended object, and the live broadcast verbal skill information at least comprises live broadcast explanation information corresponding to the to-be-recommended object; the live broadcast explanation information corresponding to the to-be-recommended object is determined according to the object description information of the to-be-recommended object. According to the embodiment of the invention, the to-be-recommended object information corresponding to the target live broadcast account and the live broadcast verbal skill information are displayed on the live broadcast plan page, and intelligent generation of the live broadcast plan is supported, so that the preparation work before live broadcast is started is simplified, and the efficiency of the preparation work before live broadcast is started is improved.
Owner:BEIJING YOUZHUJU NETWORK TECH CO LTD

Object counting method and device for enhancing text guidance by utilizing frequency characteristics

The invention belongs to the technical field of computer images, and provides an object counting method and device for enhancing text guidance by utilizing frequency characteristics. The method comprises the following steps: inputting a query image to a visual encoder to obtain visual features; inputting the object description text to a text encoder to obtain text embedding; fusing the visual features and the text features by using a fusion network to obtain fusion features; the fusion features are input into a decoder to obtain a density estimation graph of the object, the decoder comprises K cascaded adaptive frequency selection modules and one output layer, and the kth adaptive frequency selection module comprises a first convolution unit, an adaptive frequency selector and a first up-sampling unit; the kth self-adaptive frequency selector carries out filtering processing on the amplitude spectrum and the phase spectrum of the output characteristic of the first convolution unit in the frequency domain, and converts a filtering processing result into a spatial domain; according to the method, related frequency components are dynamically emphasized in decoding, the space domain and frequency domain features are combined, and accurate object positioning and accurate counting are achieved.
Owner:CHONGQING UNIV

Information pushing method and device, computer equipment and storage medium

The invention relates to an information pushing method and device, computer equipment, a storage medium and a computer program product. The method comprises the steps of obtaining candidate product features and object description features; the candidate product features and the object description features are combined to obtain shared features, at least two autocorrelation degrees of the shared features are calculated, the shared features are transformed based on the at least two autocorrelation degrees, and autocorrelation features corresponding to the at least two autocorrelation degrees are obtained; taking the candidate product features and the object description features as input features, and fusing the input features with the self-correlation features corresponding to the at least two self-correlation degrees in sequence to obtain target fusion features; and calculating an interaction degree of the to-be-pushed object to the product information based on the target fusion feature, and pushing the product information of the candidate product to a terminal of the to-be-pushed object when the interaction degree meets a preset pushing condition. By adopting the method, the information pushing accuracy can be improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Industrial robot self-adaptive grabbing method and system based on multi-mode perception

The invention provides an industrial robot self-adaptive grabbing method based on multi-modal sensing, which comprises the following steps of: 1, acquiring multi-modal data of an object and an environment through a multi-modal sensing module; the multi-mode sensing module comprises at least two sensing modules of a visual sensing unit, a touch sensing unit, a force sensing unit and an auditory sensing unit; 2, preprocessing and feature extraction are carried out on the multi-modal data, feature information of different modals is fused through a multi-modal fusion algorithm, and comprehensive object description information is generated; 3, generating an optimal grabbing strategy based on the comprehensive object description information; wherein the grabbing strategy comprises the position and posture of a grabbing point, a grabbing path and grabbing force; and fourthly, the industrial robot executes the optimal grabbing strategy, the first step to the third step are repeated, and the grabbing strategy is adjusted in a self-adaptive mode till grabbing is completed. Stable and efficient grabbing can be achieved.
Owner:HUNAN INST OF INFORMATION TECH +2

Task processing method, dialogue task processing method, task planning model training method, information processing method based on task planning model and model training platform

The embodiment of the invention provides a task processing method, a dialogue task processing method, a task planning model training method, an information processing method based on a task planning model and a model training platform. The task processing method comprises the steps of obtaining task data of a target task; the task data and the object description information of the candidate processing objects are input into a task planning model, multiple processing objects corresponding to the target task are determined, and the task planning model is used for planning the multiple processing objects of the target task; processing the task data by using the plurality of processing objects to obtain processing results respectively output by the plurality of processing objects; and determining a task processing result of the target task according to the processing results output by the plurality of processing objects. The task planning model is utilized to determine the plurality of processing objects for processing the target task, so that seamless integration between the model and the processing objects with different processing capabilities is realized, the task processing capability is expanded, and the task processing is more efficient and comprehensive.
Owner:ALIBABA (CHINA) CO LTD

Methods and systems for disambiguation of referred objects for embodied agents

This disclosure addresses the unresolved problems of tackling object disambiguation task for an embodied agent. The embodiments of present disclosure provide a method and system for disambiguation of referred objects for embodied agents. With a phrase-to-graph network disclosed in the system of the present disclosure, any natural language object description indicating the object disambiguation task can be converted into a semantic graph representation. This not only provides a formal representation of the referred object and object instances but also helps to find an ambiguity in disambiguating the referred object using a real-time multi-view aggregation algorithm. The real-time multi-view aggregation algorithm processes multiple observations from an environment and finds the unique instances of the referred object. The method of the present disclosure demonstrates significant improvement in qualifying ambiguity detection with accurate, context-specific information so that it is sufficient for a user to come up with a reply towards disambiguation.
Owner:TATA CONSULTANCY SERVICES LTD

Safety monitoring multi-modal model reasoning method and device

The invention relates to the technical field of visual reasoning, and provides a safety monitoring multi-modal model reasoning method and device. According to a user problem and a user image in a security monitoring scene, an object position in a visual scene is converted into text information, and the text information, the user problem and the corresponding user image serve as input information; obtaining visual features according to the input information through a cross-modal semantic converter; constructing a visual scene of the user image into hierarchical description comprising scene description and object description; modeling the context of the user image according to the user question, and generating a text prompt of a visual scene; reasoning is carried out through a large language model according to the visual features, the hierarchical description and the text prompt, reasoning output is obtained, and the problems that in a multi-modal scene, the complex scene perception ability is insufficient, and the large language model reasoning ability is insufficient in utilization in the prior art are solved.
Owner:709TH RESEARCH INSTITUTE CHINA STATE SHIPBUILDING CORP LTD

Interaction method and device, electronic equipment and storage medium

The invention relates to an interaction method and device, electronic equipment and a storage medium, and the method comprises the steps: receiving a first reference image, and displaying the identification information of the received first reference image in a guide language input region; the number of the first reference images is multiple; receiving a text cue word, and displaying the text cue word in the guide word input area; the text cue word comprises an object description word group, and the object description word group corresponds to the first reference image; the object description phrases are used for describing objects in the first reference image corresponding to the object description phrases; the identification information of the first reference image is interspersed in the text cue word; the number of characters between the position of the object description phrase in the text cue word and the position of the identification information of the first reference image corresponding to the object description phrase in the text cue word is less than a preset number of characters; a first video is generated based on the content in the guide language input area. The demand that a user hopes to generate a video based on a plurality of first reference images can be met.
Owner:BEIJING ZITIAO NETWORK TECH CO LTD

Distributed operation-oriented multi-agent collaborative recommendation method and system

The invention discloses a distributed operation-oriented multi-agent collaborative recommendation method and system, and relates to the technical field of intelligent recommendation, and the method comprises the steps: configuring a private domain operation agent in each independent private domain platform, deploying a big language model-based construction demand analysis agent at a user side to receive a natural language demand description input by a user, and constructing a multi-agent collaborative recommendation system; generating recommendation task instructions of different private domain platforms; the method comprises the following steps: receiving user preference description and recommendation object description, sending the description to a corresponding private domain operation agent, calling a recommendation model to generate a recommendation result list based on the received user preference description and recommendation object description, collecting recommendation result lists returned by all private domain operation agents by a demand analysis agent, and integrating and classifying the recommendation result lists to generate a comprehensive recommendation list. According to the method, an efficient, safe and extensible distributed recommendation architecture is constructed by introducing a user demand analysis agent, a private domain operation agent and a federal learning mechanism.
Owner:广东省华南技术转移中心有限公司 +1

Data information search method and device based on large model enhancement, equipment and medium

The invention discloses a data information search method and device based on large model enhancement, equipment and a medium. The method comprises the following steps: acquiring search text information input by a search object; and according to the search text information, the object description information corresponding to the search object and a target text optimization model obtained by pre-training, determining search key information of a corresponding point of the search text information. And determining search vector information corresponding to the search key information, determining a vector search result in the data index database, and determining an entity search result in the data index database according to the search key information. And determining a target search result corresponding to the search text information according to a target result evaluation model corresponding to the search object. According to the method, intention recognition, user input enhancement and better understanding of the user intention are carried out by utilizing a large model during search, search optimization is carried out in a mixed retrieval mode, and more accurate and personalized intelligent retrieval is provided in combination with a target result evaluation model.
Owner:SHANDONG JINGBEI FINANCIAL TECH CO LTD

Cross-platform page code generation method based on large model

The invention discloses a cross-platform page code generation method based on a large model, and relates to the technical field of front-end development, and the method comprises the following steps: cooperatively analyzing a UI design drawing through a plurality of special large models, and carrying out real-time cross validation and dynamic compensation on an analysis result by adopting a multi-model mutual verification mechanism; based on a predefined standard specification, dynamic adjustment is carried out through a platform adaptation rule base; generating a page object description tree embedded with the input / output processing function; cross-platform page codes are generated through a code generation engine, and sandbox testing and automatic correction are executed; outputting a target platform code by utilizing a unified compiling engine; an adaptive optimization loop is constructed based on test feedback. According to the method, through integration of multi-model collaborative analysis, platform rule dynamic adaptation, sandbox verification and closed-loop optimization, the problems of large analysis deviation, poor cross-platform compatibility, uncontrollable code quality and the like are solved, and end-to-end high-quality automatic generation from a design drawing to multi-platform codes is realized.
Owner:CHENG DU ZHONG KE JI YUN RUAN JIAN YOU XIAN GONG SI

Digital Content Creation With Dynamic Targeting

Methods, systems, and apparatus, including computer-readable storage media for generating model-generated digital content from prompts built using a combination of a base object description and targeting parameters for an intended audience. The digital content, once generated, can be served to a target audience indicated by the targeting parameters. A system implementing the methods described herein can generate content for various different audiences, indicated by different combinations of targeting parameters available on a campaign management platform serving the content. When the content is no longer being served the system can cause the digital content to be deleted or otherwise discarded. Instead of storing the content, the system can save the prompt and re-process the prompt through the model to re-generate the content. The system can further index prompts for later querying, so that the system can avoid generating new prompts over using stored prompts for content generation.
Owner:GOOGLE LLC

Video question and answer method and system based on multi-level alignment

The invention relates to a video question answering method and system based on multi-level alignment. The method comprises the following steps: generating multi-level visual features including global frame visual features and local frame visual features; generating multi-level text description according to the multi-level visual features, wherein the multi-level text description comprises global object description and local object description; and according to the generated multi-level visual features and the multi-level text description, establishing alignment between a visual mode and a text mode at an object level, a frame level and a video level, training a language model, and performing video question and answer by using the trained language model. According to the method, alignment between visual and text modes is established among object-level, frame-level and video-level multi-mode information, advanced performance can be obtained even if a small pre-training data set and few trainable parameters are adopted, and the method has wide practical value and application scenes.
Owner:INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES

An intelligent SQL generation system based on multi-level intent recognition and a generation method thereof

The application discloses a kind of intelligent SQL generation method and system based on multistage intention recognition, comprising: receiving user natural language query;Vector matching step: vector matching is carried out based on knowledge base, obtain Top-K candidate query object, and knowledge base contains multi-layer structure description information;Large language model intention understanding step: determine final query object in combination with candidate query object description;Load branch workflow, parse field and table structure, dynamically configure database table and field information;SQL generation step: generate SQL sentence in accordance with syntax rule, including automatically adding field, time condition conversion, field priority matching operation;Execute SQL sentence and return result, and record log if it fails to prompt correction.The application provides an efficient, accurate and user-friendly database query solution by natural language processing technology combined with database knowledge base.
Owner:SICHUAN ZHONGLI JIAHUA INFORMATION TECH CO LTD

Print text content display method and device based on cross-platform consistency

PendingCN121722337ADigital output to print unitsImage resolutionFont rasterization
The invention discloses a printed text content display method and device based on cross-platform consistency, and the method comprises the steps: obtaining a to-be-printed character string and a defined text object description protocol, and obtaining a text attribute structure according to the defined text object description protocol; obtaining the resolution ratio of the ink-jet printer nozzle, and enabling the resolution ratio of the PC upper computer to be consistent with the resolution ratio of the ink-jet printer nozzle according to the resolution ratio of the ink-jet printer nozzle; adopting a text shaping engine to obtain corresponding font list information according to the text attribute structure; adopting a font rasterization engine to obtain corresponding font contour information according to the font list information and the text attribute structure; obtaining the pixel width and height of the whole line of text according to the font list information and the font contour information; according to the obtained font list information, font contour information and resolution parameters, drawing to a display screen by adopting a Skia rendering engine; therefore, cross-platform measurement and drawing consistency of the text content of the ink-jet printer is realized.
Owner:SOJET MARKING TECH (XIAMEN) CO LTD

Digital content creation with dynamic targeting

Methods, systems, and apparatus, including computer-readable storage media for generating model-generated digital content from prompts built using a combination of a base object description and targeting parameters for an intended audience. The digital content, once generated, can be served to a target audience indicated by the targeting parameters. A system implementing the methods described herein can generate content for various different audiences, indicated by different combinations of targeting parameters available on a campaign management platform serving the content. When the content is no longer being served the system can cause the digital content to be deleted or otherwise discarded. Instead of storing the content, the system can save the prompt and re-process the prompt through the model to re-generate the content. The system can further index prompts for later querying, so that the system can avoid generating new prompts over using stored prompts for content generation.
Owner:GOOGLE LLC

Cochlea object description method and system, equipment, storage medium and program product

The embodiment of the invention provides a cochlea object description method and system, equipment, a storage medium and a program product, and the method comprises the steps: extracting a region-of-interest image of a cochlea object from a medical image comprising the cochlea object, carrying out the segmentation of the cochlea object, obtaining a segmentation result of the cochlea object, and carrying out the segmentation of the cochlea object according to a plurality of preset radiomics feature types, respectively extracting corresponding radiomics features from the region-of-interest image, obtaining voxels of the cochlea object from the segmentation result, calculating respectively corresponding morphological parameter values based on voxel coordinates of the voxels of the cochlea object according to a plurality of preset morphological parameters, and determining the cochlea object according to the plurality of preset morphological parameters and the plurality of morphological parameter values based on the plurality of radiomics features and the plurality of morphological parameter values. A description data set of the cochlea object is constructed to describe the cochlea object using the plurality of radiomics features and the plurality of morphological parameter values. According to the scheme provided by the embodiment of the invention, the comprehensiveness and accuracy of description aiming at the cochlea are improved.
Owner:BEIJING FRIENDSHIP HOSPITAL CAPITAL MEDICAL UNIV +1

Symptom information determination method and apparatus, electronic device, and storage medium

The application relates to a symptom information determination method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring object description information to be processed; using a target symptom recognition network to obtain corresponding target symptom information by taking the object description information as input, wherein the target symptom recognition network is obtained by machine learning training of multiple sample pairs and adjusting parameters of a preset network during the training process, each sample pair indicates a pair of object description information samples and symptom information samples, the training process comprises learning a first type correlation degree and a second type correlation degree, the first type correlation degree represents a correlation degree between two heterogeneous samples in the sample pair, and the second type correlation degree represents a correlation degree between two heterogeneous samples from different sample pairs. The application improves the accuracy and efficiency of symptom recognition. The embodiments of the application can be applied to various scenes such as cloud technology, artificial intelligence, intelligent transportation and auxiliary driving.
Owner:腾讯医疗健康(深圳)有限公司

Pixel based with object based decision making approach for driving

A method of a pixel based with object based decision making for driving, the method includes receiving, at a first machine learning process of an artificial intelligence agent, a sensed information unit; receiving, at a second machine learning process of the artificial intelligence agent, object descriptive information regarding an object captured in the sensed information unit; generating, by the first machine learning process, a pixel-based path planning output related to a suggested pixel-based path segment of a vehicle; generating, by the second machine learning process, an object-based path planning output related to a suggested object-based path segment of the vehicle; and generating, by at least in part processing the pixel-based path planning output in correspondence with the object-based path planning output, a driving related output with respect to the vehicle.
Owner:AUTOBRAINS TECH LTD

Image generation method, training method of reward model

Embodiments of the present application provide an image generation method, a reward model training method, an electronic device, a storage medium and a computer program product. The method comprises: obtaining object description text of a target object, original object image and scene information matched with a to-be-generated image carried by an image generation instruction; taking the scene information, the object description text and the original object image as input information, calling a pre-tuned visual language model to generate background image design description text adapted to the scene information, wherein when the scene information is a weak signal in the input information, the background image design description text can reflect scene preference characteristics matched with the scene information; calling a preset image generation model to generate an image of the target object matching the scene information under the constraint of the background image design description text, the image of the target object in the original object image and layout information. The image generated by the method is more matched with the scene preference corresponding to the scene information.
Owner:HANGZHOU ALIBABA INT NETWORK TECH CO LTD