Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

54 results about "Text enhancement" patented technology

Medical report generation method, model training method, equipment and medium

The invention discloses a medical report generation method, a model training method, equipment and a medium, and the model training method comprises the steps: constructing a medical report generation model framework which comprises a global semantic collaborative multi-modal enhancement module, a visual encoder, a text encoder, a medical insight analyzer and an LLM decoder; wherein the global semantic collaborative multi-modal enhancement module respectively enhances a medical image and a medical report by utilizing a selected image enhancement strategy and a text enhancement strategy, and the medical insight analyzer comprises a fine-grained structure learning device and a global context guide learning device which are connected in sequence so as to enhance the cross-modal alignment capability; and performing intelligent collaborative optimization by taking a strategy set formed by an image enhancement strategy and a text enhancement strategy and architecture configuration parameters of the medical insight analyzer as optimization targets to obtain an optimal medical report generation model. The medical report generation performance can be effectively improved.
Owner:CENT SOUTH UNIV

Multi-relation extraction error propagation optimization method and device based on data collaborative enhancement

The invention discloses a multi-relation extraction error propagation optimization method and device based on data collaborative enhancement, and the method comprises the steps: generating extended training data through grammar recombination and adversarial samples, carrying out the positioning and weighted sampling of a low-frequency relation according to an entity pair relative position, and obtaining an upper relation data set; and training an upper relation classifier based on the upper relation data set, and performing prediction through the trained upper relation classifier to obtain a prediction result and confidence distribution thereof. According to the method, through text enhancement and a weighted sampling strategy for relative position positioning based on entities, the problem of unbalanced data distribution is effectively relieved, the modeling capability of the upper relation classifier for long-tail distribution is remarkably improved, and therefore systematic optimization of the upper relation classifier for low-frequency relation recognition accuracy is achieved.
Owner:WUHAN UNIV OF SCI & TECH

Multi-modal sentiment analysis method based on text enhancement and modal completion perception fusion

The invention relates to a multi-modal sentiment analysis method based on text enhancement and modal completion perception fusion, and belongs to the field of natural language processing. The method comprises the following steps: 1, extracting text semantic features by utilizing a pre-training language model, and performing linear transformation on non-text features to form unified multi-modal input representation; 2, injecting a cross-modal enhancement module into the pre-training language model, fusing non-text information by taking a text as a core, and performing multi-modal input representation; 3, introducing a modal completion module, generating a missing modal completion representation through reconstruction loss and random modal discarding, and suppressing noise features in combination with a convolution gating structure; and guiding full-connection network learning modal weight to realize joint optimization of sentiment classification and regression. According to the method provided by the invention, on the premise of keeping the dominance of the text, through cross-modal interaction and dynamic weighting and in combination with a modal completion mechanism, the model has relatively strong emotion recognition capability.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Retrieval enhancement generation method, device and equipment based on LSH and storage medium thereof

The invention belongs to the technical field of artificial intelligence, and relates to an LSH-based retrieval enhancement generation method and device, equipment and a storage medium thereof. Carrying out serialization processing; sampling processing is carried out through the vector space model, and high-dimensional embedded vector representation is generated; inputting the high-dimensional embedded vector representation into a vector database, and carrying out matching calculation; carrying out binary code valuation conversion on a matching calculation result by adopting an LSH algorithm to obtain binary code values corresponding to all the retrieved texts after conversion; according to the binary code value, an expected retrieved text is screened out; combining and sorting the retrieval input text and the expected retrieved text to generate a to-be-enhanced text; and inputting the text to be enhanced into the text enhancement generator to generate an enhanced retrieval feedback text. The method is applied to webpage search, question and answer systems and recommendation system scenes in the field of financial services or medical services, and a faster and more accurate retrieval feedback result is provided for retrieval personnel.
Owner:PING AN TECH (SHENZHEN) CO LTD

Long-tail recommendation method and system based on multi-agent collaborative reasoning

The invention relates to a long-tail recommendation method and system based on multi-agent collaborative reasoning, and belongs to the technical field of personalized recommendation. The method comprises the following steps: performing text enhancement on user and project data, and constructing a semantic data set; constructing and training a candidate item retrieval model with long-tail perception capability; searching candidate items for a target user, and constructing a deep user portrait containing a personalized novelty exploration coefficient; starting a multi-agent collaborative reasoning process, and rearranging candidate items through multi-dimensional evaluation, strategy sorting and adaptive causal depolarization; and finally, generating a personalized recommendation list, and performing iterative optimization based on user feedback. According to the method, the recommendation task is decomposed into a collaborative reasoning step of a plurality of agents, and a personalized and causal-driven prejudice removing mechanism is introduced, so that the accuracy of overall recommendation is not damaged while the exposure and recommendation diversity of the long-tail project are improved, and a more fair and transparent recommendation service conforming to the real exploration intention of the user is provided for the user.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Text video retrieval method of fine-grained relation learning network based on energy perception

The invention provides a text video retrieval method of a fine-grained relation learning network based on energy perception. The method comprises the following steps: giving a query text and a video clip; inputting the query text into a text encoder of the CLIP, inputting the video clip into a text encoder image encoder of the CLIP, and extracting to obtain text embedding and frame embedding; inputting the text embedding and the frame embedding into a fine-grained relation learning network for text enhancement operation to obtain enhanced text embedding; taking the enhanced text embedding as a frame fusion condition to carry out frame fusion operation on the frame embedding to obtain video embedding; calculating the similarity of a text-video pair formed by the query text and the video embedding based on a cosine similarity function; and selecting the text-video pair with the highest similarity as a retrieval output result of text video retrieval. According to the method, the problem of randomness of the random text of single sampling is solved, so that semantic information of text coding is better expanded, and the final retrieval effect is improved.
Owner:SUN YAT SEN UNIVERSITY SHENZHEN +1

Cross-domain face anti-spoofing detection method and device based on multi-modal text enhancement

The application discloses a cross-domain face anti-counterfeiting detection method and device based on multi-modal text enhancement, and relates to the technical field of network information security.The method comprises the following steps: inputting two types of description texts into a pre-trained text encoder to extract text category features representing authenticity / fraud, and inputting an image into a pre-trained visual encoder to extract visual features; adding trainable text prompts to each layer of the text encoder, and adding trainable visual prompts to each layer of the visual encoder, wherein the visual prompt of each layer of the visual encoder is obtained by converting the text prompt of the current layer through a full connection layer; embedding a PFT module and a TIM module into the middle layer of each layer of the text encoder and the visual encoder to realize feature interaction and fusion, obtain the cosine similarity and the mask between the text category features and the visual features, and perform face true / false classification.The application simultaneously completes the modal feature interaction in the feature extraction process based on the PFT module and the TIM module, and improves the cross-domain detection performance.
Owner:XIAMEN UNIV

Intelligent repairing method for handwriting file fonts

The invention relates to an intelligent restoration method for handwriting archive fonts, and the method comprises the steps: decomposing a handwriting archive target image, obtaining channel images corresponding to all color channels, for each channel image, determining a character enhancement coefficient of each pixel point in a character region based on the image feature value of each pixel point in the character region of the channel image, and restoring the character enhancement coefficient of each pixel point in the character region. And then, based on each character enhancement coefficient, correcting a preset Gaussian convolution scale parameter to obtain a target Gaussian convolution scale parameter corresponding to each pixel point in each channel image, and based on a Retinex algorithm and each target Gaussian convolution scale parameter, carrying out character enhancement processing on the corresponding channel image. According to the method, channel images subjected to character enhancement processing are obtained, the channel images subjected to character enhancement processing are combined, the character enhancement image after the target image of the handwritten file is restored is obtained, and restoration of the target image of the handwritten file is achieved.
Owner:LIAONING QIDIAN EDUCATION TECH CO LTD

Input disturbance robustness improvement method of power generation type large language model

The invention relates to an input disturbance robustness improving method for a power generation type large language model, which comprises the following steps of: randomly extracting corpus data from an initial corpus to carry out first disturbance enhancement, and generating a disturbance enhancement corpus; corpora are randomly extracted from the new corpus, text enhancement is carried out in a synonym replacement and / or random character enhancement mode, and text enhancement corpora are generated; combining the initial corpus, the disturbance enhancement corpus and the text enhancement corpus to generate a training corpus; inputting data in the training corpus as a training sample into the generative large language model, randomly discarding elements of model neurons by an input layer and a hidden layer of the model, and setting the elements to zero; and the consistency loss of the generative large language model is optimized. The method has the advantages that the disturbance robustness training method based on disturbance enhancement, text enhancement and feature generalization is adopted, the anti-input disturbance capacity is enhanced, and the generalization capacity and robustness of the power generation type large language model are effectively improved.
Owner:FUJIAN YIRONG INFORMATION TECH +1

Information generation method and device based on large language model, storage medium and refrigeration equipment system

The invention discloses an information generation method based on a large language model, and belongs to the technical field of artificial intelligence, and the method comprises the steps: obtaining user input data and a task processing type; according to the task processing type, the multi-modal large model is used for processing user input data, the multi-modal large model comprises a data fusion module, an encoder module, an enhancement module, an image-text description module and a result generation module, and the data fusion module is used for receiving the user input data and fusing the user input data into text data; the encoder module is used for encoding the data output by the data fusion module, the enhancement module is used for performing text enhancement on the data output by the data fusion module according to external knowledge, and the image-text description module is used for generating image-text description data according to the data output by the encoder module; and the result generation module is used for generating information according to the data output by the encoder module, the data output by the enhancement module and the image-text description data, so that the information generation accuracy and content richness of the model can be improved.
Owner:QINDAO HAIER REFRIGERATOR CO LTD +1

A human behavior recognition method and system based on visual semantic text enhancement

PendingCN122654825AHuman bodyHuman behavior
The application relates to a human body behavior recognition method and system based on visual semantic text enhancement, and the method comprises the following steps: S10, acquiring human body behavior WIFI sensing data and human body behavior video frame data; S20, respectively extracting WIFI amplitude energy images and high-dimensional semantic feature vectors from the human body behavior WIFI sensing data and the human body behavior video frame data; S30, performing deep correlation fusion on the WIFI amplitude energy images and the high-dimensional semantic feature vectors, and calculating the probability distribution of human body action categories. The human body behavior recognition method based on visual semantic text enhancement provided by the application realizes effective fusion of semantic information and WiFi features by cross-modal similarity retrieval, and calculates the probability distribution of human body action categories based on the global feature vector obtained after fusion, so that the recognition precision of similar actions is effectively improved.
Owner:INNER MONGOLIA UNIV OF SCI & TECH

Information generation method and device based on large language model, storage medium and refrigeration equipment system

The invention discloses an information generation method and device based on a large language model, a storage medium and a refrigeration equipment system, and belongs to the technical field of artificial intelligence. According to the task processing type, the multi-modal generation type large model is used for processing the user input data, the multi-modal generation type large model comprises a data fusion module, an encoder module, an enhancement module and a result generation module, the data fusion module is used for receiving the user input data and fusing the user input data into text data; the encoder module is used for encoding the data output by the data fusion module, and the enhancement module is used for performing text enhancement according to external knowledge, the data with context features output by the encoder module and the data output by the data fusion module. And the result generation module is used for generating information according to the data output by the encoder module and the data output by the enhancement module, so that the information generation accuracy and content richness of the model can be improved.
Owner:QINDAO HAIER REFRIGERATOR CO LTD +1

Image paper-like processing method and device and electronic equipment

The invention relates to the technical field of computers, and discloses an image paper-like processing method and device and electronic equipment. The method comprises the following steps: acquiring two continuous frames of images corresponding to a target interface, wherein the target interface is an interface which contains text content and is about to realize a paper-like effect; performing text screening according to the two continuous frames of images to obtain a text region to be subjected to text enhancement; performing character stroke segmentation processing on the text region to obtain a stroke segmentation result; obtaining a stroke filling value based on the stroke segmentation result; obtaining an initial paper-like picture corresponding to the target interface; and performing enhancement processing on strokes in the initial paper-like picture by adopting the stroke filling value to obtain a final paper-like picture corresponding to the target interface. According to the application, non-native applications, pictures and video contents can achieve a paper-like effect, and the reading experience of an electronic screen is improved.
Owner:GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1

Railway signal fault text enhancement method, system and equipment based on improved EDA

The invention discloses a railway signal fault text enhancement method, system and device based on improved EDA, and the method comprises the following steps: constructing a railway field entity dictionary, carrying out the proper noun recognition of a preprocessed to-be-enhanced railway signal fault text, and outputting a structured text; on the basis of a railway signal fault related text data set, a similar word bank of railway signal fault text related vocabularies is generated by adopting a Word2Vec model and used for replacing non-professional vocabularies in railway signal fault text statements, and enhancement operation is performed on the basis of a structured text and the similar word bank of the railway signal fault text related vocabularies by adopting an EDA improvement method. Enhancing the railway signal fault text to be enhanced; and screening the enhanced railway signal fault text. According to the method, the integrity and accuracy of domain terms are ensured through an entity fixing mechanism, an enhanced sample conforming to a real scene is generated in combination with a semantic constraint editing strategy, and the problem of unbalanced data distribution is effectively solved.
Owner:YANSHAN UNIV

A robust visual question answering model training method based on contrastive learning

This invention presents a robust visual question-answering model training method based on contrastive learning, belonging to the interdisciplinary field of natural language processing and computer vision. In image enhancement, this invention employs a visual context perturbation-based enhancement method. It filters visual contexts with weak relevance to the question by analyzing the attention distribution of objects in the image and adds perturbations to these contexts to construct new image representations, allowing the model to learn visual context-independent image representations. For general question types, a strategy of removing interrogative auxiliary verbs is used for text enhancement; for other question types, a rewriting strategy is employed. Positive samples are constructed using these data enhancement methods, and then optimized using contrastive learning to learn an unbiased multimodal representation of the input information. This invention is applicable to fields such as artificial intelligence and natural language processing, enhancing model robustness and improving the accuracy of question answering on data with different scenarios or distributions.
Owner:BEIJING INST OF TECH

Text emotional state recognition method based on MEDTM model

The invention relates to a text emotional state recognition method based on an MEDTM model, which is used for researching and efficiently utilizing texts in different answers to carry out depression recognition. Firstly, a text feature extraction model is established, various text features are extracted respectively, then different text features are enhanced based on a text enhancement model, emotional state recognition is carried out through an attention fusion module and a self-supervision loss module in an MEDTM model, the problem that text emotion recognition is too short can be effectively solved, and the emotion recognition efficiency is improved. Therefore, the emotion state recognition accuracy is improved.
Owner:NORTHEASTERN UNIV CHINA

Emotion support robot man-machine interaction method based on multi-modal emotion recognition

An emotion support robot man-machine interaction method based on multi-modal emotion recognition belongs to the field of man-machine interaction technology, artificial intelligence and emotion calculation, and comprises the following steps: step 1, carrying out multi-modal emotion recognition on a user; 2, performing personalized emotion support dialogue generation based on an emotion recognition result in the step 1; and 3, transmitting the emotion recognition result in the step 1 and the emotion support dialogue in the step 2 to an interaction robot, determining a corresponding interaction action by the interaction robot according to the emotion recognition result, and outputting a personalized emotion support dialogue to a user through a TTS technology. According to the method, the multi-modal emotion data set for the college student group is constructed, the multi-modal emotion recognition model based on text enhancement is adopted, various complex emotions possibly possessed by the user at the same time can be recognized, quantitative calculation is conducted on the complex emotions, and therefore the current emotion state of the user can be recognized more accurately and more robustly, and the user experience is improved. And the limitation of single-mode identification is overcome.
Owner:NANJING FORESTRY UNIV

Text enhancement system, method and equipment for multi-scale context fusion and medium

The invention discloses a multi-scale context fusion text enhancement system, method and device and a medium, and relates to the technical field of natural language process.The method comprises the steps that cleaning processing, word segmentation processing and vectorization processing are conducted on a to-be-processed text, and an initial vector is obtained; performing position coding processing on the initial vector based on a hierarchical fusion rotation position coding mechanism to obtain a feature vector with position information; performing fusion processing on the feature vectors with the position information through a local-global attention fusion layer to obtain a multi-scale feature map; processing the multi-scale feature map based on a memory enhancement network and a hierarchical semantic aggregation network to obtain enhanced context representation; and predicting a task prediction result corresponding to the to-be-processed text based on the enhanced context representation. According to the method, the training parameter quantity and the tuning cost are remarkably reduced, and an end-to-end efficient, accurate and high-generalization-ability text understanding framework is integrally realized.
Owner:DIGITAL HEALTH CHINA TECHNOLOGIES CO LTD

Text and mask multi-mode guided dynamic video emotion recognition method

The invention belongs to the technical field of animal behavior recognition and emotion calculation, and particularly relates to a text and mask multi-modal guided dynamic video emotion recognition method, which comprises the following steps of: constructing a dynamic golden snub monkey emotion data set, and constructing multi-modal input containing videos, key points, masks and text description; obtaining a golden snub monkey video sequence, and extracting face key points and mask information; inputting the video sequence, the key points, the face mask and the text description into a double-flow guide module to generate a double-flow guide prompt; inputting a double-flow guidance prompt into a space attention priori module, and performing feature enhancement through a channel attention mechanism and a multi-head space-time perceptron; extracting features from the text description through a text enhancement network; and finally, adaptively adjusting the contribution weight of the text semantics to the visual features through a dynamic gating fusion module to realize accurate recognition of the emotion of the golden snub monkey. According to the method, the problem of cross-species dynamic emotion recognition is effectively solved through multi-modal information collaborative fusion and dynamic modeling.
Owner:NORTHWEST UNIV

Text image driven three-dimensional face generation and expression editing method and system

The application belongs to the field of computer vision processing, and provides a three-dimensional face generation and expression editing method and system based on a text picture driving, key description information is extracted from a source face text description, a control vector is generated based on the key description information and the source face picture, the source face picture is enhanced by using the control vector, and a source face picture after text enhancement is obtained; face model parameters are extracted based on the source face picture after text enhancement, a rough shape is generated by taking the face model parameters as a guide, the rough shape is enhanced in details and is rendered by mapping to generate a final three-dimensional face model; a source face picture after expression migration is generated according to a target face picture and the source face picture, the source face picture after expression migration is enhanced by using a target face text description, a source face picture after expression migration and text enhancement is obtained, and a source three-dimensional face model after expression editing is obtained by performing three-dimensional reconstruction on the source face picture after expression migration and text enhancement.
Owner:SHANDONG UNIV OF FINANCE & ECONOMICS

Text enhancement method, electronic device, storage medium

The application relates to the technical field of artificial intelligence, in particular to a text enhancement method, an electronic device and a storage medium. In the text enhancement method, original text information is acquired first, text segmentation processing is performed on the original text information to obtain an original text field, the original text field is subjected to deletion and modification processing via a language processing model to obtain a target text field, and the language processing model is obtained by optimizing and training a basic language model. Further, target text information is generated according to the target text field obtained after the deletion and modification processing, and the original text information and the target text information are integrated to form enhanced text information. In the text enhancement method, the original text field is subjected to deletion and modification processing via the language processing model to obtain the target text field, and then the original text information and the target text information are integrated to form the enhanced text information, so that the quality of sample data is improved in the text enhancement process.
Owner:CHINA PING AN LIFE INSURANCE CO LTD

Micro-blog theme mining method based on text enhancement and high-order conversation relationship orientation

PendingCN122414172AAlgorithmWeb tables
This invention discloses a microblog topic mining method based on text enhancement and high-order conversational relationship guidance, including: (1) constructing a user-level dialogue network containing high-order conversational relationships; (2) user embedding based on text enhancement: using a text enhancement method based on a large language model to interpret colloquial words to utilize their semantic information, and using a pooling function based on a multilayer perceptron to learn the mutual influence between different words, generating user representations based on text sequences; (3) self-fusion network representation: capturing the nonlinear relationship between text content and high-order network structure, and introducing an attention mechanism to model the influence of different users on the topic in the sequence, obtaining user sequence representations; (4) topic generation based on neural variational reasoning: using the user sequence representation with structural information supplementation as input to neural variational reasoning, learning potential generation factors through network reconstruction, and finally generating topics with better consistency.
Owner:TIANJIN UNIV

Visual text multi-modal action recognition method based on deep learning

The invention relates to the technical field of action recognition, in particular to a visual text multi-modal action recognition method based on deep learning, which comprises the following steps: acquiring video data and label set data, and processing the video data by adopting a video encoder; visual tokens to be fused are screened out from the label set data, and the fusion score of the tokens is calculated based on a multi-criterion token fusion strategy; performing visual token fusion based on a bipartite matching strategy and a multi-criterion fusion score; encoding the video feature input, and interacting the fused visual token and the information token to obtain a video code; performing text enhancement on the information token by adopting a text encoder to obtain text representation; cross-modal alignment is performed on the video coding and the text representation. The accuracy of the action recognition method is comprehensively improved through the multi-criterion token fusion strategy, the bipartite matching strategy, various lightweight adapters and other designs.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Pedestrian text retrieval method based on cross-modal alignment

The invention relates to a pedestrian text retrieval method based on cross-modal alignment, which comprises the following steps of: firstly, respectively extracting features from an image and a text by using a CLIP dual encoder, projecting the features to a public projection space to calculate global similarity, secondly, designing a dual-modal local prototype module to respectively store local features of the image and the text, further enhancing the fine granularity of the local features, and finally, extracting the feature of the image and the feature of the text. An independent memory bank is established for each prototype module to record local features of the prototype module, so that the model can retain feature differences in modals and realize cross-modal local alignment at the same time, and then a text enhancement module is designed to randomly mask texts and replace the texts with error words; and predicting the semantic information perception capability of the original text enhancement model through a multi-layer perceptron.
Owner:SHENYANG INST OF AUTOMATION - CHINESE ACAD OF SCI

Text automatic labeling method, system, device and terminal for judgment documents

The application belongs to the technical field of data labeling, and discloses a method, system, device and terminal for automatically labeling text about judgment documents, wherein the text with punctuation symbols is divided into sentences and then input into a Jieba Chinese parser for word segmentation; based on the word frequency word segmentation result, a dynamic programming method is used to find the path with the maximum probability, and the text word segmentation terms are stored in an intermediate database; the manually labeled text is trained through a machine learning model, and automatic labeling of the text is realized through a constructed corpus labeling library; after the annotation data in the database are scored, the labeling is automatically reloaded, and data sorting is performed according to the labeling score. Through the method combining text enhancement and semi-supervised learning, the target case document is subjected to event extraction and labeling according to preset event extraction rules, and combined with an external online database, the labor cost of text entity labeling is greatly reduced, and the efficiency and accuracy of text labeling are improved.
Owner:湖南工商大学

A pedestrian re-identification method and device based on a multi-modal adapter and a medium

The present application belongs to the technical field of data processing, and particularly relates to a pedestrian re-identification method based on a multi-modal adapter, a device and a medium. The method comprises the following steps: collecting pedestrian re-identification data, including image data of pedestrians and corresponding text descriptions; performing data enhancement on the pedestrian re-identification data, the data enhancement being used for processing image data lacking text descriptions to generate text descriptions corresponding to the image data; performing text enhancement on the text descriptions based on a multi-modal large language model; constructing a multi-modal adapter to generate text adaptation embedding and image adaptation embedding, and superimposing the text adaptation embedding and the image adaptation embedding with original text embedding and image embedding to obtain text fusion features and image fusion features; and performing pedestrian re-identification prediction by using the text fusion features and the image fusion features to realize pedestrian re-identification. The present application constructs a multi-modal adapter, has fewer trainable parameters, enhances the adaptation capability of the model to a target domain, and significantly reduces the calculation efficiency.
Owner:CENT SOUTH UNIV

Sample set optimization method and apparatus, device, medium, product

The application relates to a sample set optimization method and device, equipment, medium and product. The method comprises the following steps: obtaining an original sample set, including a plurality of training samples and corresponding supervision labels; determining the influence degree of each training sample in the original sample set after the supervision label is changed according to an influence function, removing part of the training samples with relatively high influence degrees, and obtaining a balanced sample set; implementing text enhancement processing based on part of the training samples in the balanced sample set, expanding the training samples through text enhancement, and obtaining an augmented sample set; removing the outlying training samples in the augmented sample set based on the clustering results of the deep semantic information of each training sample in the augmented sample set, and obtaining an optimized sample set. The optimized sample set obtained by the application is rich in sample quantity and excellent in sample quality, is suitable for training a deep learning model corresponding to a related downstream task, makes the trained deep learning model more easily convergent, and can obtain a higher prediction accuracy.
Owner:GUANGZHOU HUANJU SHIDAI INFORMATION TECH CO LTD

Non-portrait image quality enhancement method and device based on multi-region adaptation

This invention relates to the field of image processing, specifically disclosing a method and system for enhancing the image quality of non-human images based on multi-region adaptive processing. The method includes: acquiring a non-human image to be enhanced; dividing the non-human image into text regions, subject regions, and non-subject regions, where the subject regions and non-subject regions correspond to different enhancement requirements; generating a text stitching image based on the content of the text regions of the non-human image, enhancing the text stitching image using a preset text enhancement model, and generating a text region enhancement result based on the model output; enhancing the subject regions and non-subject regions respectively using corresponding enhancement models according to their different enhancement requirements; and finally, fusing the three enhancement results to obtain the enhanced result. Thus, through differentiated processing by region, different enhancements of text, subject, and non-subject can be achieved simultaneously in the same image, improving the enhancement effect and exhibiting good versatility and scalability.
Owner:GUANGZHOU GUANGZHUIYUAN INFORMATION TECH CO LTD

High-efficiency fine tuning method for large language model for fine-grained text classification

The invention discloses a large language model efficient fine tuning method for fine-grained text classification, and relates to the technical field of natural language processing and machine learning application. Comprising the following steps: step 1, performing data construction and enhancement: acquiring a text data set with fine-grained category labels, and performing data enhancement by adopting a combined text enhancement strategy, step 2, performing fine adjustment of multi-strategy fusion, and step 3, performing performance verification and confusion analysis: evaluating the overall classification performance of the model on the verification set, wherein the overall classification performance comprises the accuracy rate, the precision rate, the recall rate and the F1 score.
Owner:INSPUR QILU SOFTWARE IND

Causal event prediction system and training method thereof

The invention belongs to the related technical field of machine learning, and discloses a causal event prediction system and a training method thereof, the system comprises a graph reweighting unit, a text enhancement unit, an alignment unit, a fusion unit and a prediction unit, the graph reweighting unit fuses text semantics into a graph structure by fusing event semantic similarity and a graph adjacency matrix, and the prediction unit is used for predicting a causal event. The method comprises the following steps of: acquiring a graph representation of an event node by using a graph reweighting unit to obtain a graph representation of the event node, fusing a graph topology constraint in text representation learning by using a text enhancement unit through structural context injection, graph flattening and a sparse attention mechanism to obtain a text representation of the event node, and realizing deep fusion of text semantics and graph topology information through the graph reweighting unit and the text enhancement unit. And then the two representation spaces are aligned through a comparison unit, bimodal information is integrated in combination with a fusion mechanism and then is input into a prediction unit, a prediction result of the to-be-predicted event is obtained, semantic rationality and causal coherence of the prediction result are further ensured, and the performance of causal graph event prediction is remarkably improved.
Owner:HUAZHONG UNIV OF SCI & TECH