Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

1258 results about "Image content" patented technology

Multi-modal named entity recognition method based on semantic alignment and cross-modal graph fusion

The invention belongs to the technical field of natural language processing and multi-modal information extraction, and particularly relates to a multi-modal named entity recognition method based on semantic alignment and cross-modal graph fusion, which comprises the following steps: S1, acquiring a data sample containing a text sequence and image content; s2, encoding the text and the image into vectors respectively; s3, similarity is calculated through a trainable bilinear function, and optimization is carried out through loss comparison; s4, cross-modal attention is used to enhance association information between modals; s5, determining the proportion of reserved image information through a modal matching module; s6, introducing a gating mechanism to dynamically fuse visual and text features; s7, realizing local and global information complementation by a cross-modal graph fusion model; and S8, inputting the fused representation into the CRF layer to predict the entity type. According to the method, fine semantic alignment can be realized in a weak image-text correlation context, and balance between local entity recognition and global semantic understanding can be achieved.
Owner:ANHUI UNIVERSITY OF TECHNOLOGY

Cross-modal eye fundus image generation method and system based on generative adversarial network

The invention discloses a cross-modal eye fundus image generation method and system based on a generative adversarial network, relates to the technical field of medical image processing, and constructs an eye fundus focus perception and edge consistency generative adversarial network by taking a cyclic consistency generative adversarial network as a baseline. The core of the method is that a lesion perception mixed attention module is embedded in a bottleneck layer of a generator so as to strengthen the extraction capability of fine features of a lesion area; an edge information extraction module is designed, and key edge features are accurately extracted in combination with Roberts edge detection, wavelet transform and non-local mean denoising; and a joint loss function containing edge consistency loss is constructed, and the semantic consistency of a focus structure during cross-modal generation is ensured by minimizing the feature difference between the source image and the generated image. According to the method, the problems of disordered content, inconsistent structure and unstable training of the generated image in the prior art are effectively solved, and the simulation degree and clinical availability of the generated image are remarkably improved.
Owner:SUZHOU UNIV

Potential safety hazard real-time identification method and system based on multi-modal large model

The invention belongs to the technical field of data processing, and particularly relates to a potential safety hazard real-time identification method and system based on a multi-modal large model, the system comprises the multi-modal large model, a field knowledge base and a multi-modal inference engine, the multi-modal large model is responsible for visual feature extraction and scene semantic understanding of an input field image, and the field knowledge base is responsible for field knowledge base analysis; processing a natural language query provided by a user; the domain knowledge base stores construction safety related laws and regulations, guidelines and historical cases, and factual basis is provided for the system through a structured storage and efficient retrieval mechanism; the multi-modal reasoning engine coordinates the whole process of visual understanding, task decomposition, knowledge retrieval and report generation, and is a core control module for realizing multi-modal reasoning and decision making. And refined understanding of entities and hidden dangers in a construction scene is realized. According to the system and the method thereof, the image content and the text specification can be dynamically fused, and missing detection or misjudgment caused by modal splitting in a traditional method is avoided.
Owner:CEC ANSHI (CHENGDU) TECH CO LTD

Document interpretation and report generation method and device, equipment and medium

The invention relates to the technical field of natural language processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a document interpretation and report generation method, device, equipment and medium, which comprises the following steps: receiving an original document set to generate a structured document object, executing optical character recognition on an image content set to generate a recognition text set, the recognition text set and the text content set are combined into a unified text sequence, element item extraction is executed based on the interpretation template parameter set to generate an interpretation element set, a retrieval enhancement context is retrieved and generated from the domain knowledge base, and the unified text sequence, the interpretation template parameter set and the retrieval enhancement context are input into a language model to generate an interpretation result. And generating report content based on the historical report template set. According to the method, automatic closed loop of document interpretation and report generation is realized through multi-modal unified processing and semantic enhanced reasoning, the efficiency is improved, and the manual dependence and compliance risk are reduced.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Intelligent paper marking system and method for talent selection and recruitment

The invention relates to the technical field of human resource management, and provides an intelligent paper marking system and method for talent selection and recruitment. The method comprises the following steps: cutting test paper according to a preset test paper template by using an image processing technology to obtain a plurality of plates corresponding to the test paper; converting the image content of each section of the cut test paper into a text format which can be edited and processed by adopting a character recognition technology; setting a preliminary scoring standard according to question type information and knowledge point information corresponding to the test paper, and optimizing the preliminary scoring standard by using a large language model; performing semantic understanding and logic analysis on the converted test paper text answers from a plurality of scoring dimensions including semantic accuracy, logic integrity and knowledge point coverage by using two preset scoring large models to give corresponding scoring information, and marking advantage information and defect information in the answers; and obtaining a comprehensive score corresponding to the test paper according to score information printed by each score big model for each question of each test paper.
Owner:SHENZHEN TALENT GROUP CO LTD

Backlight source dynamic refreshing method and system based on content change frequency

The invention provides a backlight dynamic refreshing method and system based on content change frequency, and is applied to the field of image data processing. According to the method, the inter-frame variable quantity of the current image content of the display screen is pre-collected, whether the image refreshing speed is matched with the backlight response or not is judged, the pixel change rate in unit time is calculated based on the frame difference method, the content change frequency index is reasonably divided, the backlight PWM parameters are dynamically adjusted in combination with user setting, and the display quality is improved. The response coordination of the display screen under the dynamic image is effectively improved, when the frame rate fluctuation is recognized, the gray scale compensation frame is introduced and the brightness change slope is limited, and the problems of smear, splash screen and the like caused by asynchronous refreshing can be relieved.
Owner:SHENZHEN GAOXINXING TECH CO LTD

Intelligent teaching-assistant question-answering system with enhanced multi-modal knowledge graph

The invention belongs to the technical field of artificial intelligence and educational informatization, and relates to a multi-mode knowledge graph enhanced intelligent teaching-assistant question-answering system. According to the method, the knowledge graph construction technology, the multi-modal content analysis technology and the large language model reasoning enhancement technology are comprehensively applied, and the semantic understanding, knowledge integration and reasoning generation capabilities of the intelligent teaching assisting system in an education and teaching scene are improved. The related technology comprises layout analysis of textbook documents, semantic description generation of image content, entity and relation extraction of text content, multi-modal knowledge graph construction and question and answer reasoning and natural language generation combined with the knowledge graph. Through cooperative application of the technologies, semantic interconnection can be carried out on various modal information such as texts, images and tables in the textbook, and a searchable and traceable textbook-level knowledge network is formed.
Owner:NORTHEASTERN UNIV CHINA

Agent-based image analysis and processing system and method, and storage medium

The present invention relates to the technical field of computer image processing, and in particular to an agent-based image analysis and processing method and system, and a storage medium. Image quality analysis and estimation is performed by means of a pre-trained image quality analysis model, and a corresponding processing strategy and a corresponding image processing sequence are generated, thereby achieving automatic image quality analysis and evaluation and facilitating selection of different processing strategies based on different image content, so as to satisfy processing requirements of various application scenarios. A corresponding image processing model is retrieved on the basis of the result of matching between the processing strategy and basic information of each image processing model in a knowledge base, and then called according to the image processing sequence, and an automatic image processing flow control model automatically controls the image processing model to execute image processing according to the image processing sequence and performs monitoring, thereby achieving automatic image processing, reducing the requirement of manual intervention, and saving the time and cost.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Sorting control system of flow cytometry sorting instrument

PendingCN121113842AIndividual particle analysisCell sorterControl system
According to the sorting control system of the flow cytometry sorting instrument, liquid drop images are shot through a camera, and liquid drop states are quantitatively evaluated and monitored on the basis of image content information. In order to determine the liquid drop breaking state, a liquid drop image stroboscopic shooting mode is designed to observe the liquid drop state, and the STM32H7 provides a synchronous control signal. When the droplet generation state is stable, the droplet breaking image is close to a stationary state. The stroboscope lamp is controlled to be turned on at different moments in a period, so that different states of liquid drop generation can be observed. The time difference between the detection time and the charging time is quantitatively determined through liquid drop delay control, and accurate delay control over sorting is achieved. And finally, a full-automatic algorithm is adopted to calibrate the liquid drop time delay, the liquid drop time delay is continuously tried to be changed, and when the shooting brightness of the sorted CCD camera under a certain liquid drop time delay is the highest, the time delay is considered to be the current liquid drop time delay, so that full-automatic calibration of the liquid drop time delay is realized.
Owner:SUZHOU INST OF BIOMEDICAL ENG & TECH CHINESE ACADEMY OF SCI

Video generation method

The invention belongs to the technical field of artificial intelligence, and particularly relates to a video generation method, a video generation device, a computer readable medium, electronic equipment and a computer program product. The method comprises the steps that a reference video and a reference image are acquired, the video content of the reference video comprises the posture change process of a first human body in a motion state, and the image content of the reference image comprises the static display effect of a second human body in a specified view angle; performing feature extraction on the reference video to obtain global attitude features and local attitude features; and generating a target video according to the global posture feature, the local posture feature and the reference image, wherein the video content of the target video comprises the posture change process of the second human body in the motion state. The motion detail expressive force generated by the video can be improved, and the global and local posture coordination of the human body can be enhanced.
Owner:WEBANK (CHINA)

Method and system for intelligently and dynamically adjusting brightness of LED backlight source

The invention relates to the technical field of LED light source adjustment, and discloses an intelligent dynamic adjustment method and system for the brightness of an LED backlight source. The method comprises the following steps: acquiring an image frame sequence and real-time brightness; a brightness distribution histogram and inter-frame brightness variation are extracted based on the image frame sequence, image content change feature extraction is carried out, and an image monitoring result is generated; calculating target brightness and adjustment parameters of backlight partitions according to a monitoring result, performing deviation analysis and frequency amplitude weight matching in combination with real-time brightness, and generating partition adjustment demand data; smoothing the demand data to obtain a partition brightness curve and a smooth real-time index; calculating an intermediate compensation coefficient by correlating the fusion curve and the index to obtain a compensation coefficient set and a curve driving basis; judging the matching degree of the compensation demand and the amplitude weight to obtain a matching evaluation result; and finally, adjusting the brightness of the partitions, compensating inter-frame brightness variation, and determining a display output state. According to the method, the brightness adjustment precision and response speed under the dynamic content are improved.
Owner:SHENZHEN YOUZHI ELECTRONICS CO LTD

Dynamic mesh coding with simplified topology

A mesh decoder reconstructs geometry information of a dynamic mesh from a coded mesh bitstream of the dynamic mesh. The reconstructed geometry information include data specifying vertices of the dynamic mesh. The decoder also reconstructs connectivity information of the dynamic mesh which includes data specifying faces of the dynamic mesh. The decoder refines the reconstructed connectivity information based on the reconstructed geometry information to generate refined connectivity information. The decoder further reconstructs an attribute image of the dynamic mesh from the coded mesh bitstream which includes image content to be applied to faces of the dynamic mesh. The decoder refines the reconstructed attribute image based on the reconstructed geometry information to generate refined attribute image. Based on the reconstructed geometry information, the refined connectivity information, and the refined attribute image, the decoder reconstructs the dynamic mesh which can be rendered for display.
Owner:GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD

Multi-mode vertical field detection method based on multi-agent cooperation

The invention discloses a multi-mode site detection method based on multi-agent cooperation, and relates to the field of site detection, and the method comprises the steps: obtaining social media text content and social media image content issued by a target user for a target topic; according to the social media text content and the social media image content, utilizing a multi-agent collaborative multi-mode site detection framework to determine a site of the target user for the target topic; according to the multi-agent collaborative multi-modal site detection framework, firstly, a multi-modal feature analysis layer is deployed, and nine types of professional agents collaboratively implement content-context-emotion three-dimensional feature extraction; secondly, a double-wheel standing debate layer is constructed, and multi-view argumentation and retro are achieved through structured double-wheel adversarial debate; and finally, logic verification and evidence evaluation are executed through a verification decision-making layer, so that inference chain integrity and result reliability are ensured. According to the invention, the accuracy of multi-mode vertical field detection is improved.
Owner:PEOPLES POLICE UNIV OF CHINA (INT LAW ENFORCEMENT COOP INST OF THE MINISTRY OF PUBLIC SECURITY CHINA PEACEKEEPING POLICE TRAINING CENT)

Airborne visible light image automatic splicing method

The invention discloses an airborne visible light image automatic splicing method, and relates to the field of unmanned aerial vehicle image processing, and the method comprises the steps: obtaining a plurality of images collected in the flight process of an unmanned aerial vehicle, and the corresponding spatial position information and attitude information; on the basis of the spatial position relation between the images and the image overlapping information, determining image pairs capable of being registered, and constructing a connection map with the images as nodes and the image pairs capable of being registered as edges; detecting whether spatial connection fracture caused by image missing exists in the atlas or not, if so, inserting a virtual node and constructing a virtual image comprising an edge region and a transition region; and further calculating geometric transformation parameters between the images, resampling all image contents to a unified coordinate system, and executing pixel-level fusion processing in an image overlapping region. According to the invention, the problem of discontinuous image splicing of the unmanned aerial vehicle due to the existence of a no-fly zone, a privacy protection zone and the like of the unmanned aerial vehicle is solved.
Owner:DI RUI TIANCHENG INFORMATION TECH (BEIJING) CO LTD

Dual-machine target positioning method based on semantic prior and spatial intersection constraint

The invention discloses a dual-machine target positioning method based on semantic prior and spatial intersection constraint. The method comprises the following steps: 1, acquiring images of a camera 1 and a camera 2, and determining the spatial orientation of a camera coordinate system relative to a reference world coordinate system; step 2, obtaining a target normalized coordinate under a coordinate system of the camera 1; step 3, obtaining a target normalized coordinate under a coordinate system of the camera 2; 4, converting the normalized coordinates of the target into geometric rays in a three-dimensional space; 5, constructing an optimization function with the minimum distance sum of squares; step 6, performing semantic-level foreground and background division on the image content through a YOLOv5 network, and constructing a three-dimensional space feature point library of a semantic background; and 7, performing single-view semantic neighborhood retrieval and depth compensation by using the three-dimensional space feature point library of the semantic background constructed in the step 6, and realizing target three-dimensional positioning in a high-altitude long-distance complex scene. According to the invention, the detection robustness and the positioning precision of the high-altitude long-distance target in the environment are improved.
Owner:XIDIAN UNIV

Generated image quality evaluation method based on pre-training cross-modal feature alignment embedding

The invention relates to a generated image quality evaluation method based on pre-training cross-modal feature alignment embedding, and compared with the prior art, the method solves the defect of limited quality evaluation precision caused by neglecting generation prompt information and neglecting local view quality degradation in the traditional image quality evaluation process. The method comprises the following steps: acquiring and generating an image quality evaluation data set; performing image multi-granularity decomposition; constructing a dynamic quality prototype; constructing a quality evaluation model; training a quality evaluation model; and obtaining a quality evaluation result. In a cross-modal feature alignment space constructed by a pre-training large model, object display index region division of an image is completed, dynamic quality prototype fitting data quality manifold distribution is constructed, an enhanced expression statement is constructed in combination with text prompt information when the image is generated, the consistency inspection of image content and text prompt is completed, and the image content and text prompt consistency is improved. And meanwhile, global view quality evaluation is considered, and the image quality evaluation task precision is remarkably improved in multiple scenes.
Owner:ANHUI UNIV

Fine-grained clothing image retrieval method and device based on large language model common sense knowledge injection

The application discloses a fine-grained clothing image retrieval method and device for injecting common sense knowledge of a large language model. The method first extracts fine-grained visual features of an input image through an image encoder and optimizes them in combination with a low-rank adapter to improve the representation ability at the image patch level. Then, a pre-trained large language model generates attribute-enhanced common sense knowledge context to enrich the image attribute representation, thereby helping the model understand and reason about unknown attribute information in an open scenario. The application introduces a switchable modal prompt and interpolation mechanism to ensure that the agent embedding can be dynamically supplemented when attributes or text are missing. During the retrieval process, through an attribute-guided cross-modal attention mechanism, fine-grained image content matching is performed based on the relationship between image features and attribute-enhanced context. Through multi-modal feature alignment and optimization, the application improves the accuracy and robustness of clothing image retrieval in an open world scenario.
Owner:ZHEJIANG GONGSHANG UNIVERSITY

Multi-mode intelligent semantic understanding and abstract generation system and method based on HDMI (High Definition Multimedia Interface) stream

The invention provides a multi-mode intelligent semantic comprehension and abstract generation system and method based on an HDMI stream, and the system comprises an HDMI input module which is used for receiving a video signal outputted by external equipment; the image content streaming analysis module is configured to perform content segmentation, optical character recognition, layout structure extraction and graphic element detection on the video signals and output structured image data; the audio acquisition module is used for acquiring an audio signal and preprocessing the audio signal; the voice recognition module is used for transferring the preprocessed audio signal into a voice text; and the multi-modal fusion and semantic understanding module is configured to fuse the structured image data and the voice text, generate summary information and output the summary information to the result output module. According to the invention, real-time synchronous analysis of the multi-modal content and intelligent generation of the abstract are realized.
Owner:BEIJING YUNJIANXIN TECH CO LTD

Vehicle display system, vehicle display method, and storage medium

A vehicle display system, by executing a computer-readable program, being configured to: display multiple image contents in separate display areas of a display device arranged in a vehicle; detect a visual field range of an occupant of the vehicle; determine whether a frame rate of a certain image content displayed in the visual field range satisfies a lowest reference, the certain image content being one of the multiple image contents; and in response to determining that the frame rate of the certain image content does not satisfy the lowest reference, reduce a processing load of a different process while maintaining a display process of the certain image content in the visual field range of the occupant.
Owner:DENSO CORP

Multi-modal knowledge graph completion method for metal material data

The invention relates to the technical field of knowledge graph completion, in particular to a metal material data-oriented multi-modal knowledge graph completion method. The method comprises the following steps: acquiring multi-modal data of a target data source, and preprocessing the multi-modal data to obtain a text entity, a text entity vector and an image content vector; extracting a text entity relationship triple and a table entity relationship triple from the text entity and the preprocessed structured table data to obtain an entity relationship triple; updating the text entities based on the text entity vector and the image content vector to obtain an entity set, and disambiguating the entities to obtain a standardized entity set; and constructing a target knowledge graph according to the standardized entity set and the entity relationship triad, and performing relationship completion on the target knowledge graph. In this way, the problems that an existing method lacks semantic alignment and information complementation technologies among modals, and complex dependency relationships among entities in the multi-modal knowledge graph cannot be fully considered can be solved.
Owner:辽宁材料实验室

Dark field detail dynamic enhancement method and system of LED backlight source

The invention relates to the technical field of image enhancement, in particular to a dark field detail dynamic enhancement method and system for an LED backlight source, and the method comprises the following steps: setting a brightness threshold T1 and a brightness threshold T2 according to the brightness component of an input image frame, the dynamic backlight range of the associated LED backlight source and a gamma curve; according to the method, the pixel brightness is subjected to fine-grained division through partition judgment of the brightness component of the input image frame, so that more accurate dark field area positioning is realized, and the local linear transformation coefficient is constructed for the dark field mask area by using the original brightness component, so that the subsequent enhancement operation has continuity and edge retention characteristics; and adaptive classification of the image content is realized through a structure complexity threshold value, then differential enhancement strategies such as multi-direction texture enhancement, normal direction sharpening and curved surface smoothing are extracted respectively, and gain distribution conforming to different region characteristics is established, so that the dark field detail identification granularity is higher.
Owner:HONGBAO FURUI TECHNOLOGY (SHENZHEN) CO LTD

Portable digital video camera configured for remote image acquisition control and viewing

A wearable digital video camera (10) is equipped with wireless connection protocol and global navigation and location positioning system technology to provide remote image acquisition control and viewing. The Bluetooth® packet-based open wireless technology standard protocol (400) is preferred for use in providing control signals or streaming data to the digital video camera and for accessing image content stored on or streaming from the digital video camera. The GPS technology (402) is preferred for use in tracking of the location of the digital video camera as it records image information. A rotating mount (300) with a locking member (330) on the camera housing (22) allows adjustment of the pointing angle of the wearable digital video camera when it is attached to a mounting surface.
Owner:CONTOUR IP HLDG LLC

Personalized image content generation and optimization method and system fusing generative AI

The invention provides a personalized image content generation and optimization method and system fusing generative AI, and relates to the technical field of AI. The personalized image content generation and optimization method comprises the steps of analyzing a historical interaction track and a current intention expression of a target user based on a deep perception network, and constructing a multi-dimensional behavior portrait; user groups with similar generation preferences are identified by using a group feature recursive quantization technology, and a group wisdom feature map is constructed; performing multi-level mapping analysis on the user features and the group wisdom feature map, anchoring the group affiliation relationship of the target user and concretizing the guide features; constructing a multi-level guide vector according to current intention expression and guide features, performing accurate mapping on hidden space expression of an image generation model, quantizing deviation in real time and triggering intelligent calibration; and reconstructing a group distribution structure based on the evaluation data of the target user to realize dynamic evolution of the features. According to the invention, personalized demands of users can be accurately grasped, and the accuracy of image generation and the satisfaction degree of the users are improved.
Owner:SMIC WANYE TECHNOLOGY CO LTD

Medical image artifact recognition and elimination method based on big data technology

The invention discloses a medical image artifact identification and elimination method based on a big data technology, and relates to the technical field of medical image processing, and the method comprises the steps: carrying out the preprocessing of collected image data based on gray normalization, and then carrying out the semantic segmentation and ROI positioning of the image content; and performing lesion segmentation on the positioned image content, selecting lesion features for quantification and fusion, performing model verification, and performing distributed deployment on the verified lesion segmentation model. According to the method, the problems of high missed detection rate of small nodules and high missed diagnosis risk of malignant lesions are solved through the lesion detection model, the false positive rate is reduced, the recall rate of the malignant lesions is improved, and through the lesion segmentation model, the segmentation adaptability to lesions of different sizes is improved, clear segmentation boundaries are obtained, surgical planning is assisted, and boundary positioning errors are reduced.
Owner:眉山市人民医院 +1

Stereoscopic vision optimization method and system for naked-eye 3D large screen

The invention discloses a stereoscopic vision optimization method and system for a naked-eye 3D large screen, and particularly relates to the technical field of naked-eye 3D vision optimizing.The method comprises the steps that environment illumination and audience positions are sensed in real time through multi-sensor fusion, and an environment light field model is established; glare crosstalk noise is predicted based on physical simulation, and self-adaptive suppression and compensation are carried out in combination with human eye visual sensitivity and image content features; a virtual camera is dynamically generated according to the real-time positions of the eyes of the audience, and a lightweight neural radiation field renderer is used for real-time re-rendering, so that motion parallax is realized; and finally, intelligently fusing the glare compensation layer and the perspective correction layer, coding and outputting to a screen. The system correspondingly comprises an environment perception module, a glare compensation module, a perspective rendering module and a fusion coding module. The naked-eye 3D large-screen display method effectively inhibits ambient light interference, improves the quality and immersion of a stereoscopic picture under different visual angles, and is suitable for naked-eye 3D large-screen display under outdoor and complex illumination environments.
Owner:ANHUI SHENGZI TECH CO LTD

Image meaning analysis scene consistency evaluation system based on visual model

The invention relates to the technical field of image processing, discloses an image meaning analysis scene consistency evaluation system based on a visual model, and aims to solve the problem of insufficient recognition stability in complex environments such as illumination variation, angle deviation and local shielding in the prior art. The system comprises an image input module used for receiving and preprocessing a plurality of images; the visual large model analysis module is used for carrying out multi-dimensional semantic feature extraction on the preprocessed image; the structured description generation module is used for converting the semantic features into structured text description in a unified format; the scene consistency evaluation module is used for performing logic consistency analysis on the structured text descriptions of the plurality of images; and the result output module is used for generating and outputting a final consistency evaluation report. According to the technical scheme, the method can effectively improve the recognition stability of the system in a complex environment, achieves the multi-dimensional semantic understanding of the image content, and remarkably reduces the misjudgment rate of multi-view image evaluation.
Owner:SHANGHAI SHANHAO INTELLIGENT TECH DEV CO LTD

Image scene conversion method and device, equipment and storage medium

The invention discloses an image scene conversion method and device, equipment and a storage medium. Firstly, an image to be processed and a conversion cue word used for indicating an image scene conversion direction and a specific attribute adjustment requirement can be obtained, the conversion cue word is input into a pre-trained scene noise optimizer, and structured guide noise accurately matched with the cue word is generated through noise optimization processing and used as target noise. Performing vector encoding on the to-be-processed image and the conversion prompt word through an image encoder and a text encoder to obtain an image feature vector and a text feature vector, and performing feature fusion on the image feature vector and the text feature vector by using a cross attention mechanism to generate a fusion vector; and finally, jointly inputting the fusion vector and the target noise into a pre-trained scene conversion model, and performing conditional decoding generation processing to obtain a scene conversion image corresponding to the to-be-processed image. According to the invention, disordered disturbance caused by noise injection is avoided, and drift of the image content in the conversion process is effectively inhibited.
Owner:HANGZHOU WANGDAO HLDG CO LTD

RAG-based pdf intelligent retrieval and generation method and system

The application discloses a kind of PDF intelligent retrieval and generation method and system based on RAG, by obtaining the document data of input, using the classification model established in advance to parse document data, extract text content and image content to form first data set;Using deep learning model to the image content in first data set carries out feature extraction, while the text content in first data set applies natural language processing technology to carry out semantic analysis, obtains multimodal feature set;According to multimodal feature set, application information integration algorithm is uniformly encoded and is handled to generate second data set, if detecting the integrity of fusion feature vector in second data set is lower than preset threshold value, then supplementary context semantic analysis fills in missing information;Using preset index construction mechanism to the clustering processing of fusion feature vector in second data set, generates the retrieval index library containing classification index structure.The application improves the accuracy and comprehensiveness of document retrieval.
Owner:HUNAN ZHIXUE YOUKE INFORMATION TECHNOLOGY CO LTD +1

Parsimonious inference on convolutional neural networks

The disclosed system incorporates a new learning module, the Learning Kernel Activation Module (LKAM), at least serving the purpose of enforcing the utilization of less convolutional kernels by learning kernel activation rules and by actually controlling the engagement of various computing elements: The exemplary module activates / deactivates a sub-set of filtering kernels, groups of kernels, or groups of full connected neurons, during the inference phase, on-the-fly for every input image depending on the input image content and the learned activation rules.
Owner:IRIDA LABS

Image processing method and apparatus, device, medium, and program product

Disclosed in embodiments of the present application are an image processing method and apparatus, a device, a medium, and a program product. The method comprises: acquiring a reference image and a content description text; performing encoding and copying processing on the reference image to obtain N image encoded features to be attenuated; performing feature attenuation processing on said N image encoded features to obtain N attenuated image encoded features; and according to image content described in the content description text, performing video generation processing on the N attenuated image encoded features to obtain a target video. By applying the embodiments of the present application, it can be ensured that image content of each video frame in the target video and image content of the reference image have both content diversity and content consistency.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD