Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

779 results about "Textual information" patented technology

Multi-target commodity identification method, device and system based on multi-modal data processing

The invention relates to the technical field of intelligent vending, solves the problem that in the prior art, commodity identification cannot be accurately carried out in a multi-target scene, and provides a multi-target commodity identification method, device and system based on multi-modal data processing. The method comprises the following steps: acquiring multiple frames of real-time images in a commodity transaction scene; performing preprocessing and label information extraction on the real-time image, and determining character information corresponding to the target image and the commodity label; performing instance segmentation on the target image, and determining commodity position information; performing feature extraction on the target image, and determining commodity image feature information; according to pre-collected multi-source privatized data in an intelligent vending scene, performing fine adjustment and optimization processing on the open-source multi-modal visual language model to obtain a multi-modal large model; and inputting the commodity image feature information and the text information into the multi-modal large model for information fusion, and determining a commodity target identification result. According to the invention, commodity identification can be accurately carried out in a multi-target scene.
Owner:YOPOINT SMART RETAIL TECH LTD

Detecting and using non-textual information in human speech

PCT designated stage expiredWO2025141559A1Speech recognitionMachine learningEncoder decoderAcoustic model
Automatic recognition of non-verbal messages in speech, and in particular to detection or analysis of prosodic multilayered analysis of intonation units such as prosodic unit prototypes and their multi-labeled variations, may form a hierarchical classification for the analysis of non¬ verbal information or cues in speech. A speech captured by a microphone is fed to a weakly- supervised deep learning acoustic model for speech recognition and transcription, that may be based on encoder-decoder Transformer architecture, such as Whisper by OpenAI. The model is trained to output multiple words form the text in the captured speech, to identify Intonation Units (IUs) that include one or more words, and associate non-verbal labels to each of the IUs. The labels may indicate a prototype, a discourse function (such as a conversation action), an emotion, an emphasis, or an attitude, as well as a genre of a part of, or whole of, the entire captured speech.
Owner:YEDA RES & DEV CO LTD

Virtual reality scene three-dimensional reconstruction method and system based on multi-source data fusion

The invention discloses a virtual reality scene three-dimensional reconstruction method and system based on multi-source data fusion, and relates to the technical field of underground commercial and traffic integrated space scene three-dimensional reconstruction. The system constructs a multi-dimensional data set through synchronous acquisition of point cloud, images, poses, positioning and semantic text information; a three-dimensional convolutional neural network is adopted to establish an AI reconstruction fusion model, and a three-dimensional model is generated; the system calculates a structural integrity evaluation coefficient JGPG, a fusion consistency evaluation coefficient RHPG and a VR interaction adaptability evaluation coefficient VRJH, compares the coefficients with corresponding thresholds, triggers an optimization strategy, and finally completes model adaptability optimization and packaging output, thereby improving the reduction degree and interaction performance of a virtual reality scene.
Owner:GUANGDONG ZHONGKE KAIZER INFORMATION TECH CO LTD

Government affair service content navigation method and system based on large language model

The invention relates to the technical field of multi-modal large language models, in particular to a government affair service content navigation method and system based on a large language model, and the method comprises the following steps: receiving multi-modal information input by a user through texts, voices or pictures; the voice is converted into a text, and character information in the picture is analyzed by using an OCR (Optical Character Recognition) technology; fusing multi-modal data, and inputting the fused multi-modal data into a large language model for semantic understanding and context association analysis; matching items are retrieved in combination with a local government affair knowledge base, and an initial recommendation list is generated; the recommendation result is displayed through the intelligent assistant, and interaction optimization options are provided; the method has the beneficial effects that the current government affair service content navigation mode is optimized through the multi-modal capability, the retrieval enhancement capability and the content generation capability of the large language model, so that the navigation is more modal and more intelligent, and meanwhile, the privacy and authority of data reply are ensured.
Owner:INSPUR SOFTWARE CO LTD

Recommendation method for enhancing semantics and interest perception by using large language model

The invention discloses a recommendation method for enhancing semantics and interest perception by using a large language model. The method comprises the following steps: firstly, performing semantic modeling on unstructured text information such as user comments, article description and the like by utilizing the powerful capability of a large language model in semantic comprehension and user preference modeling aspects, so as to improve the deep perception capability of a recommendation system on user interests and article semantic attributes; then, through a semantic feature alignment and discretization strategy, the problem that continuous semantic representation generated by a large language model is incompatible with features of a traditional recommendation system in the aspect of an expression structure is solved; finally, unified modeling of semantic information and traditional recommendation signals is achieved through a recommendation integration mechanism, and recommendation performance and model interpretability are improved.
Owner:SOUTHEAST UNIV

Character recognition method and device, computer equipment and storage medium

The invention provides a character recognition method and device, computer equipment and a storage medium. According to the scheme, the slice scanning image is firstly acquired, then the label area is identified according to the slice scanning image, and the character information is close to the label, so that the determined target area is larger than the area of the label so as to cover the character information. And then the slice scanning image is intercepted according to the target area, and a target image only containing key information is obtained. And character recognition is carried out on the target image to obtain a character recognition result. And finally, the character recognition results are screened by using a rule base, and a target recognition result is screened out. According to the scheme, through tag region identification and target region interception, the character identification range is narrowed, the identification efficiency is improved, and the waste of computing resources caused by processing a large amount of irrelevant information is reduced. The rule base can better adapt to complex slice images and constantly changing requirements, information needed by a user is accurately screened out, and the character recognition efficiency is greatly improved.
Owner:MOTIC CHINA GROUP CO LTD

Knowledge-intensive visual question and answer automatic data generation method and device

The invention relates to a knowledge-intensive visual question and answer automatic data generation method and device, and the method comprises the steps: constructing an original visual data set containing the professional knowledge of a target domain according to a static image, a video stream and multimedia content; extracting a representative frame sequence, converting the audio information into text information, and extracting character information in the static image to construct a structured visual instance database; according to the prompt text meeting the preset professional depth condition, establishing a three-level prompt system containing domain knowledge, an evaluation standard and a generation specification; generating a corresponding visual question and answer pair data set according to the dynamic cooperation of the main agent and the domain expert agent; generating a multi-agent quality evaluation system according to the quality evaluation result; and designing a difficulty grading mechanism according to the negative example sample. According to the method, the professionality, the accuracy and the diversity of the visual question and answer data are remarkably improved, and reliable data support is provided for training and evaluation of a multi-modal large model.
Owner:TSINGHUA UNIVERSITY

Systems and methods for generating and using semantic images in deep learning for classification and data extraction

Disclosed is a new document processing solution that combines the powers of machine learning and deep learning and leverages the knowledge of a knowledge base. Textual information in an input image of a document can be converted to semantic information utilizing the knowledge base. A semantic image can then be generated utilizing the semantic information and geometries of the textual information. The semantic information can be coded by semantic type determined utilizing the knowledge base and positioned in the semantic image utilizing the geometries of the textual information. A region-based convolutional neural network (R-CNN) can be trained to extract regions from the semantic image utilizing the coded semantic information and the geometries. The regions can be mapped to the textual information for classification / data extraction. With semantic images, the number of samples and time needed to train the R-CNN for document processing can be significantly reduced.
Owner:CROWDSTRIKE

Stamp area character recognition method and device and nonvolatile storage medium

The invention discloses a seal area character recognition method and device and a nonvolatile storage medium. The method comprises the following steps: determining a candidate seal image area in an image according to color information of pixel points in the image; point-by-point sliding convolution processing is carried out on the candidate seal image area through a multi-scale annular convolution kernel group, so that a target seal image area is determined in the candidate seal image area, and the multi-scale annular convolution kernel group comprises a plurality of convolution kernels which are of concentric ring structures and have different radius lengths; mapping the target seal image area into a rectangular expanded image, and identifying and extracting a character image to be identified in the rectangular expanded image; and performing identification processing on the character image to be identified to obtain a seal text corresponding to the target seal image area. The technical problem that the text image processing efficiency is low due to the fact that the text information of the seal area cannot be effectively recognized in the related technology is solved.
Owner:CHINA TELECOM CORP LTD

Large vision-language model reasoning method based on multi-modal tree search

The invention provides a large vision-language model reasoning method based on multi-modal tree search, which is characterized in that multi-step multi-modal auxiliary information is called and generated through an auxiliary tool, a candidate reasoning path of a target task is simulated and evaluated in combination with a prediction expansion tree search mechanism, and an optimal path is selected through self-voting, so that the reasoning efficiency is improved. Therefore, through a vision-text interleaving reasoning framework and an extension strategy during testing, vision and text information are fully utilized, the reasoning ability of a large vision-language model in a complex multi-step reasoning task is remarkably improved, a more accurate reasoning result can be obtained, model fine adjustment is not needed, and the reasoning efficiency is improved. Therefore, the reasoning result can be obtained more quickly, and the scheme of the invention has good practical application value.
Owner:FUDAN UNIVERSITY

Artificial intelligence chatbot

Methods and systems for interacting with users via a chatbot. A natural language query is received and processed by submitting a search query to a search engine. The search engine identifies relevant information including textual information and images for formulating a response. The identified information and query are submitted to a Large Language Model which generates a response displayed via the chatbot. The response may include textual information and relevant images. The system can extract text from images of documents and convert textual information into numerical vector representations for processing. Selectable options based on clustered relevant information can be provided to users for query refinement when appropriate. The chatbot interface enables natural language interactions while leveraging search capabilities and Artificial Intelligence to provide informative and helpful responses with both text and visual elements.
Owner:HONEYWELL INTERNATIONAL INC

Big model-based receipt identification method, system and equipment

The invention belongs to the field of bill recognition, and provides a bill recognition method, system and equipment based on a large model, and the method comprises the steps: obtaining bill image data containing text information, and carrying out the preprocessing of the obtained bill image data; evaluating image quality, language distribution and layout complexity of the preprocessed image, and determining a processing path of document image data by integrating evaluation results; according to the determined processing path, calling a corresponding optical character recognition model to extract text information in the document image data; performing deep semantic analysis on the extracted text data by using a pre-trained deep learning large model, and extracting key feature information; and performing rule compliance analysis on the text information or the extracted key feature information, and generating an identification analysis report according to an analysis result. According to the invention, the processing speed and accuracy are improved, the adaptability to various document formats is enhanced, and the limitation in the prior art is effectively solved.
Owner:INSPUR GENERSOFT CO LTD

Remote cooperation method based on polymorphic large model and related equipment

The invention discloses a remote cooperation method and device based on a polymorphic large model, electronic equipment and a storage medium, and relates to the technical field of computers. The method comprises the following steps: acquiring a visual field video stream of a production field and voice data inquired by field personnel, wherein the visual field video stream is acquired by AR glasses worn by the field personnel in real time; analyzing the view video stream to obtain first record content; extracting character information in the voice data, and matching the character information with a pixel position of a target object to obtain second record content; and inputting the first record content and the second record content into a pre-trained polymorphic large model, generating multimedia summary response content, and feeding back the multimedia summary response content to a user terminal of the field personnel, and enabling the field personnel to perform field inspection according to the multimedia summary response content displayed by the user terminal. The remote cooperation efficiency can be improved.
Owner:SHENZHEN PINGFANG SCI & TECH

Method, system and software for processing text

A method for processing a piece of textual information. The piece of textual information is parsed into a set of plaintext input tokens. Each of the plaintext input tokens is individually transformed using a first binary data transformation, to achieve a set of binary input tokens. Each of the set of binary input tokens is transformed individually or collectively, using an embedding data transformation, into one or several vectorized input tokens. The one or several vectorized input tokens is / are fed to a first neural network. A response is received from the first neural network in the form of one or several vectorized output tokens.
Owner:LIVEARENA TECHNOLOGIES INC

CAD file summary information automatic extraction method and device, medium and product

The invention provides a CAD file summary information automatic extraction method and device, a medium and a product, and relates to the technical field of computer picture word process.The method comprises the steps that a target CAD file is analyzed through a CAD file analysis development tool, and first word information and geometric attribute information corresponding to the first word information are obtained; retrieving the first character information according to the retrieval keyword to obtain second character information; according to the corresponding geometric attribute information in the second text information, carrying out positioning adjustment display through a CAD file analysis development tool to obtain a picture file; calling a first large language model to perform character recognition and arrangement processing on the picture to form a text list; and calling the second large language model to obtain summary information of the target CAD file according to a preset prompt template and the text list. According to the scheme, when a large batch of CAD files are processed, the processing efficiency is remarkably improved.
Owner:SUZHOU XINKUAIZHUANG TECH CO LTD

Financial reimbursement data processing method, device and system

The invention relates to the technical field of computer data processing, and discloses a financial reimbursement data processing method, device and system.The financial reimbursement data processing method comprises the following steps that multi-modal data of a reimbursement voucher is obtained, and the multi-modal data is stored in a block chain; the multi-modal data comprises pictures, voices and videos of reimbursement vouchers; recognizing characters on the picture based on a convolutional neural network, and extracting and analyzing character information of key frames of the voice and the video by using a natural language processing model; and performing feature extraction on the character information to obtain character feature information, judging whether the character feature information falls into a reimbursement directory, and performing local model fitting according to the character feature information. According to the method, the dynamic verification rule is generated based on the historical reimbursement data, the bill anomaly probability is calculated in real time through the anomaly coefficient formula, and the anomaly detection accuracy is improved by combining the error coefficient and keyword matching degree and the historical record relevance index.
Owner:BEIJING SHIJITAN HOSPITAL CAPITAL MEDICAL UNIVERSITY

Natural language generation using knowledge graph incorporating textual summaries

Techniques are provided for producing an answer to a question regarding a domain. A natural-language textual sequence representing the question is received. From a knowledge graph associated with the domain, first and second textual passages are received using rankings corresponding to the natural-language textual sequence, a first textual summary is received summarizing textual information in a first vicinity of the first textual passage, and a second textual summary summarizing textual information in a vicinity of the second textual passage is received. An answer to the question is obtained using a language model by encoding a first intermediate output based on the natural-language textual sequence, the first textual passage, and the first textual summary, encoding a second intermediate output based on the natural language textual sequence, the second textual passage, and the second textual summary, and decoding a concatenation of the first and second intermediate outputs. An output is provided.
Owner:WRITER INC

Sign language recognition glasses device, system and method

The invention discloses a sign language recognition glasses device, system and method, and belongs to the technical field of intelligent equipment, and the sign language recognition glasses device comprises a glasses frame which is used for supporting all components of the device; the camera is arranged on the front side of the glasses frame and is used for capturing hand actions in real time to generate image information; the data processing module is embedded in the glasses frame, electrically connected to the camera and used for processing the image information and executing sign language recognition; the main control unit is arranged on an ear rod of the glasses frame, is electrically connected to the data processing module and is used for coordinating the operation of each component; the display module is arranged at the positions of the lenses of the glasses frame, electrically connected to the main control unit and used for displaying character information obtained after sign language recognition, communication between hearing-impaired people and common people is not limited by specific places any more, and instantaneity and convenience of communication are greatly improved.
Owner:HANGZHOU YIDIAN ELECTRIC TECH CO LTD

Intelligent psychological counseling system based on multi-mode large model and cognitive behavior therapy

The invention provides an intelligent psychological counseling system based on a multi-mode large model and a cognitive behavior therapy, and the system comprises the steps: recognizing voice from a user side through the multi-mode large model, and obtaining text information and emotional state information; analyzing the text information by using a large language model, carrying out conversation with the user side, and extracting psychological state information of the user side from the conversation; a cognitive behavior treatment strategy matched with the user side is determined according to the emotional state information, so that a voice response to the user side is generated according to the matched cognitive behavior treatment strategy by using a multi-mode large model; by combining a multi-modal large model and a large language model, accurate and efficient natural language understanding of a user side is realized, highly anthropomorphic voice response is provided, dependence on professional psychological consultants is reduced, limitation of psychological consultation services on time and regions can be broken through, large-scale consultation service requirements can be met in an all-weather and low-cost manner, and the psychological consultation service experience is improved. The man-machine interaction experience in the consultation process is improved, and high-quality and high-practicability psychological consultation and treatment services are provided.
Owner:BEIJING INST OF TECH

Scene graph construction method, target retrieval method and related devices

The invention discloses a scene graph construction method, a target retrieval method and a related device. The scene graph construction method comprises the following steps: constructing a topological node set and a topological edge set; for each topological node in the topological node set, taking a set of visual information, text information and spatial position information of the topological node as multi-modal data of the topological node; for each topological node in the topological node set, determining multi-modal data of an object node corresponding to the topological node and a connection edge set corresponding to the topological node based on the visual information of the topological node; and constructing a target scene graph based on the multi-modal data of each topological node, the multi-modal data of each object node, the topological edge set and the connection edge set corresponding to each topological node. The target scene graph has semantic consistency and spatial continuity, and the expression ability and cross-scene generalization performance of the scene graph in a complex environment are remarkably improved.
Owner:北京数原数字化城市研究中心

Scheme detection method and device for complex image-text mixed file

The invention provides a scheme detection method and device for a complex image-text mixed file, and the method comprises the steps: extracting standard review items and standard review contents in a standard manual, and carrying out the structural processing, and obtaining a standard review file; converting a to-be-detected file into a to-be-detected image, extracting a to-be-examined multi-scale candidate region image from the to-be-detected image by using an initial recognition model and a matching model obtained by OCR and pre-training, and extracting character information of the multi-scale candidate image to obtain character information of the candidate region; inputting the multi-scale candidate region image and the standard review file into a lightweight hybrid twin network obtained by pre-training for image feature matching, and outputting a matched feature image; inputting the matched feature image, the standard review file and the candidate area text information into a multi-modal large model obtained by pre-training for compliance analysis, and outputting a detection result; according to the method, the automation and intelligence degree of review of the complex image-text mixed file can be remarkably improved.
Owner:ZHEJIANG SHUANGYUAN TECH CO LTD

Image detection method and device and electronic equipment

The invention relates to the technical field of image processing, in particular to an image detection method and device and electronic equipment. The method comprises the steps of obtaining image information of a to-be-detected object; wherein the image information comprises one or more of text features, film features and portrait features. The text feature comprises text information. The characteristics of the covering film comprise at least one of a fitting characteristic and a flattening characteristic of the covering film and the to-be-detected object. The portrait features comprise at least one of face features and body features. And detecting whether the image information is abnormal according to a preset detection rule. And if abnormity occurs, generating alarm information. Wherein the exceptions comprise one or more of text feature detection exceptions, film feature detection exceptions and portrait feature detection exceptions. According to the scheme, the detection efficiency and accuracy are improved, and the manual quality inspection cost is saved.
Owner:AISINO CORPORATION

Maintenance guidance method and device for household appliances

The invention provides a maintenance guidance method and device for a household electrical appliance. The method comprises the following steps: constructing a unified vector database in advance based on text data, image data and audio data related to a household electrical appliance fault; acquiring field condition information of the household electrical appliance, wherein the field condition information comprises video information, audio information and / or text information; respectively carrying out image matching, audio matching and / or text matching on the field condition information based on the unified vector database, and outputting knowledge fragments as retrieval results; and inputting a retrieval result into the large language model, and outputting a maintenance guidance scheme of the household electrical appliance. According to the method, a retrieval and reasoning framework based on multi-modal feature direct fusion can be constructed, information dimension reduction loss is avoided by directly analyzing image, audio and text features of faults, and the comprehensive accuracy of fault diagnosis is improved to 95% or above; and through a dynamic knowledge retrieval and enhanced generation technology, the risk of'knowledge illusion 'possibly occurring in the professional field of the general large language model is effectively reduced.
Owner:HANGZHOU ROBAM APPLIANCES CO LTD

Display device and voice recognition method

The invention provides a display device and a voice recognition method, and the method comprises the steps: obtaining an initial annotation text in a training corpus corresponding to voice data, and recognizing characters contained in the initial annotation text; determining whether the characters are replaced according to a preset probability; under the condition that replacement is needed, searching a replacement candidate list of the to-be-replaced character from the replacement table, and randomly selecting a target character from the replacement candidate list to replace the to-be-replaced character; generating a target annotation text according to the initial annotation text, the to-be-replaced character and the target character; and training a voice recognition model according to the target labeled text, and recognizing the to-be-recognized voice according to the trained voice recognition model. According to the method, character information with the same or similar pronunciation is introduced in the speech recognition model training process, so that the model can recognize characters in speech more accurately in the recognition process, the word error rate is effectively reduced under the condition of context bias, and the problem of low recognition accuracy in the current speech recognition process is solved.
Owner:HISENSE ELECTRONIC TECH (WUHAN) CO LTD

Intelligent order generation and management system

The invention discloses an intelligent order generation and management system, and relates to the technical field of information processing. The system comprises a data acquisition module; the multi-modal information processing module is used for calling the data in the database and performing character extraction through unified analysis and structured processing; the knowledge graph construction module is used for establishing a knowledge graph of product entities and entity relationships according to the products, mapping the extracted text information to entity nodes of the knowledge graph, and performing semantic association of different source information to form preprocessed order information; and the intelligent generation module is used for generating standardized JSON structured order information from the preprocessed order information and automatically inputting the standardized JSON structured order information into a database. According to the invention, through a unified data acquisition and storage mechanism, automatic identification and collection of various information formats are realized, manual arrangement is not needed, the work processing efficiency is greatly improved, and the complexity of manual input and proofreading is reduced.
Owner:SUZHOU ZHIYOU QIUSUO INTELLIGENT TECHNOLOGY CO LTD

Method and device for synchronizing voice and mouth shape of digital human and electronic equipment

The invention relates to the field of digital people, and discloses a method and a device for synchronizing voice and mouth shape of a digital person, and electronic equipment. The method comprises the steps of obtaining natural language information, converting the natural language information into character information, determining a video frame and an audio frame based on the character information, and controlling the digital person to speak when the first frame time of the audio frame is the same as the first frame time of the video frame, so that the voice and mouth shape of the digital person can be kept consistent.
Owner:WUHAN AOTUO INTELLIGENT TECH CO LTD

Advertisement video collection generation method and device, equipment and storage medium

The invention relates to the technical field of image processing, solves the problem that in the prior art, a video collection matched with an advertisement target cannot be accurately screened and generated according to different scene types, and provides an advertisement video collection generation method and device, equipment and a storage medium. The method comprises the steps that the scene type of an advertisement to be put is acquired, and the scene type comprises an introduction scene, a product display scene, a user experience scene and a problem solving scene; according to the scene type, utilizing a preset text matching algorithm to carry out similarity evaluation on character information related to the video clip content and a preset text template, and determining a target similarity corresponding to each video clip; and according to the target similarity and a preset similarity threshold value, synthesizing the video clips of which the target similarity is greater than the similarity threshold value into a video collection. According to the method, the video collection which better conforms to the advertisement target and can accurately convey the advertisement information is generated, and the advertisement effect is improved.
Owner:BEIJING LIANSHI LEGEND NETWORK TECH CO LTD

Systems and methods to control polarization on social media platforms

Methods and systems are described for control of polarization including generation of a suggested response to a social media post. A social media post is received from a device associated with a user. A first taxonomy of the post's textual information and a second taxonomy for a connected user account are determined. The first and second taxonomies and a predetermined condition are compared. A response intended for the connected user account is generated with a third taxonomy similar to the second taxonomy based on the comparison. Related apparatuses, devices, techniques, and articles are also described.
Owner:ADEIA GUIDES INC

Natural language generation using knowledge graph incorporating textual summaries

Some embodiments relate to receiving a natural-language textual sequence representing; retrieving, from a knowledge graph, a first textual passage and a second textual passage based on rankings with respect to the natural-language textual sequence, a first textual summary summarizing textual information in a first vicinity of the first textual passage, and a second textual summary summarizing textual information in a vicinity of the second textual passage; obtaining the textual output in response to the textual input using a language model by encoding a first intermediate output based on the natural-language textual sequence, the first textual passage, and the first textual summary, encoding a second intermediate output based on the natural language textual sequence, the second textual passage, and the second textual summary, and decoding a concatenation of the first intermediate output and the second intermediate output; and providing an output to a user based on the textual output.
Owner:WRITER INC