Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

525 results about "Textual information" patented technology

Knowledge-intensive visual question and answer automatic data generation method and device

The invention relates to a knowledge-intensive visual question and answer automatic data generation method and device, and the method comprises the steps: constructing an original visual data set containing the professional knowledge of a target domain according to a static image, a video stream and multimedia content; extracting a representative frame sequence, converting the audio information into text information, and extracting character information in the static image to construct a structured visual instance database; according to the prompt text meeting the preset professional depth condition, establishing a three-level prompt system containing domain knowledge, an evaluation standard and a generation specification; generating a corresponding visual question and answer pair data set according to the dynamic cooperation of the main agent and the domain expert agent; generating a multi-agent quality evaluation system according to the quality evaluation result; and designing a difficulty grading mechanism according to the negative example sample. According to the method, the professionality, the accuracy and the diversity of the visual question and answer data are remarkably improved, and reliable data support is provided for training and evaluation of a multi-modal large model.
Owner:TSINGHUA UNIVERSITY

Stamp area character recognition method and device and nonvolatile storage medium

The invention discloses a seal area character recognition method and device and a nonvolatile storage medium. The method comprises the following steps: determining a candidate seal image area in an image according to color information of pixel points in the image; point-by-point sliding convolution processing is carried out on the candidate seal image area through a multi-scale annular convolution kernel group, so that a target seal image area is determined in the candidate seal image area, and the multi-scale annular convolution kernel group comprises a plurality of convolution kernels which are of concentric ring structures and have different radius lengths; mapping the target seal image area into a rectangular expanded image, and identifying and extracting a character image to be identified in the rectangular expanded image; and performing identification processing on the character image to be identified to obtain a seal text corresponding to the target seal image area. The technical problem that the text image processing efficiency is low due to the fact that the text information of the seal area cannot be effectively recognized in the related technology is solved.
Owner:CHINA TELECOM CORP LTD

Method, system and software for processing text

A method for processing a piece of textual information. The piece of textual information is parsed into a set of plaintext input tokens. Each of the plaintext input tokens is individually transformed using a first binary data transformation, to achieve a set of binary input tokens. Each of the set of binary input tokens is transformed individually or collectively, using an embedding data transformation, into one or several vectorized input tokens. The one or several vectorized input tokens is / are fed to a first neural network. A response is received from the first neural network in the form of one or several vectorized output tokens.
Owner:LIVEARENA TECHNOLOGIES INC

Natural language generation using knowledge graph incorporating textual summaries

Techniques are provided for producing an answer to a question regarding a domain. A natural-language textual sequence representing the question is received. From a knowledge graph associated with the domain, first and second textual passages are received using rankings corresponding to the natural-language textual sequence, a first textual summary is received summarizing textual information in a first vicinity of the first textual passage, and a second textual summary summarizing textual information in a vicinity of the second textual passage is received. An answer to the question is obtained using a language model by encoding a first intermediate output based on the natural-language textual sequence, the first textual passage, and the first textual summary, encoding a second intermediate output based on the natural language textual sequence, the second textual passage, and the second textual summary, and decoding a concatenation of the first and second intermediate outputs. An output is provided.
Owner:WRITER INC

Sign language recognition glasses device, system and method

The invention discloses a sign language recognition glasses device, system and method, and belongs to the technical field of intelligent equipment, and the sign language recognition glasses device comprises a glasses frame which is used for supporting all components of the device; the camera is arranged on the front side of the glasses frame and is used for capturing hand actions in real time to generate image information; the data processing module is embedded in the glasses frame, electrically connected to the camera and used for processing the image information and executing sign language recognition; the main control unit is arranged on an ear rod of the glasses frame, is electrically connected to the data processing module and is used for coordinating the operation of each component; the display module is arranged at the positions of the lenses of the glasses frame, electrically connected to the main control unit and used for displaying character information obtained after sign language recognition, communication between hearing-impaired people and common people is not limited by specific places any more, and instantaneity and convenience of communication are greatly improved.
Owner:HANGZHOU YIDIAN ELECTRIC TECH CO LTD

Scheme detection method and device for complex image-text mixed file

The invention provides a scheme detection method and device for a complex image-text mixed file, and the method comprises the steps: extracting standard review items and standard review contents in a standard manual, and carrying out the structural processing, and obtaining a standard review file; converting a to-be-detected file into a to-be-detected image, extracting a to-be-examined multi-scale candidate region image from the to-be-detected image by using an initial recognition model and a matching model obtained by OCR and pre-training, and extracting character information of the multi-scale candidate image to obtain character information of the candidate region; inputting the multi-scale candidate region image and the standard review file into a lightweight hybrid twin network obtained by pre-training for image feature matching, and outputting a matched feature image; inputting the matched feature image, the standard review file and the candidate area text information into a multi-modal large model obtained by pre-training for compliance analysis, and outputting a detection result; according to the method, the automation and intelligence degree of review of the complex image-text mixed file can be remarkably improved.
Owner:ZHEJIANG SHUANGYUAN TECH CO LTD

Natural language generation using knowledge graph incorporating textual summaries

Some embodiments relate to receiving a natural-language textual sequence representing; retrieving, from a knowledge graph, a first textual passage and a second textual passage based on rankings with respect to the natural-language textual sequence, a first textual summary summarizing textual information in a first vicinity of the first textual passage, and a second textual summary summarizing textual information in a vicinity of the second textual passage; obtaining the textual output in response to the textual input using a language model by encoding a first intermediate output based on the natural-language textual sequence, the first textual passage, and the first textual summary, encoding a second intermediate output based on the natural language textual sequence, the second textual passage, and the second textual summary, and decoding a concatenation of the first intermediate output and the second intermediate output; and providing an output to a user based on the textual output.
Owner:WRITER INC

News broadcasting method based on artificial intelligence and related device

The invention provides a news broadcasting method based on artificial intelligence and a related device. The method comprises the following steps: acquiring text information and chart information corresponding to target news; performing information processing on the character information and the chart information to obtain a target broadcast text and corresponding target feature information; analyzing the target feature information to obtain a reference content attribute and a reference broadcast emotion; determining a target broadcast style according to the reference content attribute and the reference broadcast emotion; obtaining a target broadcast demand of a target user; performing voice synthesis on the target broadcast text based on the target broadcast style and the target broadcast demand through a preset AI broadcast model to obtain a target broadcast voice; and in response to a voice broadcasting operation of the target user, broadcasting with the target broadcasting voice. The picture and text information of the news is processed, and the adaptive voice is synthesized through the AI broadcasting model in combination with the personalized requirements of the user, so that the news broadcasting quality is improved.
Owner:JIANGXI RONG MEDIA BRAIN TECHNOLOGY CO LTD

Systems and methods for vision-language planning (VLP) foundation models for autonomous driving

Methods and systems for training an autonomous driving system using a vision-language planning (VLP) model. Image data is obtained from a vehicle-mounted camera, encompassing details about agents situated within the external environment. Via image processing, the system identifies these agents within the environment. A Bird's Eye View (BEV) representation of the surroundings is then generated, encapsulating the spatiotemporal information linked to the vehicle and the recognized agents. Execution of the VLP machine learning model begins by extracting vision-based planning features from the BEV, and receiving or generating textual information characterizing various attributes of the vehicle within the environment. Text-based planning features are extracted from this textual information. To enhance model performance, a contrastive learning model is engaged to establish similarities between the vision-based and text-based planning features, and a predicted trajectory is output based on the similarities.
Owner:ROBERT BOSCH GMBH

CAD (Computer Aided Design) drawing watching method and device supporting cross-platform font configuration, equipment and medium

The invention provides a CAD drawing watching method and device supporting cross-platform font configuration, equipment and a medium, and belongs to the technical field of data processing. The CAD drawing reading method supporting the cross-platform font configuration comprises the following steps: receiving a drawing reading request sent by a front-end operating system; according to an image analysis algorithm, a CAD drawing corresponding to the drawing reading request is analyzed, drawing analysis information is obtained, and the drawing analysis information comprises text information; mapping fonts corresponding to the character information are selected from a built-in font library of a back-end server according to applicable operation platforms of the CAD drawing, and the built-in font library comprises mapping fonts applicable to different operation platforms respectively; and using the mapping font to render a CAD drawing, and sending a reading file corresponding to the CAD drawing to a front-end operating system. The problems of font missing, format incompatibility, rendering incompleteness and drawing information loss or deformation caused by platform incompatibility fonts in the drawing viewing process in the prior art can be solved.
Owner:INSPUR HONGQI (SHANDONG) DIGITAL TECHNOLOGY CO LTD

Logistics express bill automatic identification and bill number extraction method based on rule configuration

The invention provides a logistics express sheet automatic identification and express sheet number extraction method based on rule configuration, and relates to the technical field of data processing, and the method comprises the steps: 1, carrying out the layout structure analysis and geometric correction of an input logistics express sheet image, calculating geometric correction parameters through evaluating the deformation degree of the image and combining the structural features of the logistics express sheet, and obtaining a logistics express sheet number; performing correction processing on the logistics express sheet image based on the geometric correction parameter to obtain a standardized image; and 2, performing optical character recognition processing on the standardized image, extracting all character information in the standardized image, performing word segmentation, entity recognition and semantic understanding on the recognized text by utilizing natural language processing, obtaining a key semantic unit, and organizing recognition results into a text data set. According to the invention, automatic identification and bill number extraction of the logistics bill are realized, and the logistics information processing efficiency and accuracy are improved.
Owner:XIAMEN FINGERPRINT TECH CO LTD

Generating customized content using a generative model

Disclosed are systems and methods that generate a natural language prompt that is configured to be processed by a generative model, such as a large language model (LLM), and includes certain user information to facilitate the determination and / or generation of customized content for users of an online platform. For example, textual information associated with certain user information may be extracted and aggregated and incorporated into one or more natural language prompts, which may be processed by a generative model, such as an LLM, to generate a particular output based on the type of customized content being sought and / or generated for the user. The output may then be processed to determine and / or generate the customized content or the user.
Owner:PINTEREST INC

Cooking assisting method, device and system, electronic equipment and readable storage medium

The invention relates to a cooking assisting method, and the method comprises the steps: obtaining the input information of a to-be-cooked food material, and the input information of the to-be-cooked food material comprises at least one of a food material picture and food material text information; according to the input information of the to-be-cooked food material, recommended menu options and corresponding pin guidance are provided for a user, and the pin guidance is used for guiding the user to insert a temperature probe into the to-be-cooked food material; and after the user selects a menu option and successfully inserts the temperature probe into the food material to be cooked according to the guide of the contact pin, generating a cooking control parameter corresponding to the menu option. The cooking auxiliary method can improve the convenience of food material cooking and the success rate of one-key cooking. The invention further relates to a cooking auxiliary device and system, electronic equipment and a readable storage medium.
Owner:MAXEYE SMART TECH CO LTD

Data storage method and device

The invention provides a data storage method and device, and relates to the field of storage. The method comprises the following steps: acquiring a PDF file; analyzing the PDF file to obtain an analysis result; the analysis result comprises text information obtained by performing text extraction on the PDF file or outline information of the PDF file; associating the analysis result with the PDF file to obtain a to-be-stored file; the to-be-stored file comprises a PDF file and an analysis result; and storing the to-be-stored file. Therefore, when the unstructured data such as the PDF file is stored, the original file of the PDF file is stored, and the analysis result of the PDF file and the PDF file are associated and then are together stored in a disk. Therefore, the user can access the analysis result corresponding to the PDF file at the same time when accessing the PDF file subsequently, so that the situation that the user performs manual algorithm analysis on the PDF file to obtain required data when accessing the PDF file each time is avoided, and the workload of secondary file processing of the user is reduced.
Owner:HUAWEI TECH CO LTD

Method and system for identifying attack infrastructure

There is provided a method for identifying attack infrastructure, performed by a computing system. The method may comprise collecting network traffic data of a host device that is the target of a security incident, acquiring, from the network traffic data, first data regarding port numbers recorded in the network traffic data, second data regarding protocols recorded in the network traffic data, and third data regarding textual information included in network packets recorded in the network traffic data or category information provided by network equipment, automatically identifying attack infrastructure corresponding to the first data, the second data, and the third data by using a predefined classification model and blocking network access of an Internet Protocol (IP) address associated with the attack infrastructure, wherein the attack infrastructure corresponds to resources utilized by an attacker during the security incident.
Owner:KOREA INTERNET & SECURITY AGENCY

Electronic license verification method and device, electronic equipment, storage medium and program product

The invention discloses an electronic license verification method and device, electronic equipment, a storage medium and a program product, and belongs to the technical field of data information security processing. The method comprises the steps that information of a to-be-verified electronic certificate is acquired, and the information of the to-be-verified electronic certificate comprises an electronic certificate image corresponding to the to-be-verified electronic certificate and an OFD layout file of the to-be-verified electronic certificate; verifying the effectiveness of the OFD layout file to obtain a first verification result; under the condition that the first verification result is that verification is passed, performing image processing on the electronic license image to obtain first character information corresponding to the electronic license image; performing similarity comparison on the first character information and second character information corresponding to the OFD layout file to obtain a second verification result; and determining whether the to-be-verified electronic certificate is a real electronic certificate or not according to the second verification result. Therefore, the accuracy of verifying the electronic certificate can be improved.
Owner:CHINA MOBILE (XIONGAN) ICT CO LTD +3

Auxiliary system and method for intelligent glasses suitable for aging based on human-computer interaction

The invention discloses an auxiliary system and method for intelligent glasses suitable for aging based on human-computer interaction, and relates to the technical field of human-computer interaction, and the method comprises the steps: capturing two paths of video pictures, shooting a current picture according to a preset voice instruction, recognizing the picture information on the current picture, verifying an assistant through an identity verification mechanism, and carrying out the verification of the assistant. According to the method, two paths of video pictures and the picture information are displayed, an assistant passing identity verification adjusts a camera of the intelligent glasses according to the displayed picture and observes a specific scene to provide assistance for a glasses wearer, the two paths of video pictures are synchronously captured and subjected to differentiation processing, environment picture jitter is eliminated with the help of a picture stabilization algorithm, and the intelligent glasses are more intelligent. The key features and the key character information of the current picture are extracted and compared with the preset database, permissions are distributed to different assistant roles through identity verification, picture information display is intelligently adjusted, interaction convenience and assistance efficiency are remarkably improved, and daily requirements of old people are comprehensively met.
Owner:SHENZHEN SMART CLOUD TECHNOLOGY CO LTD

Interaction method and device based on virtual reality, equipment and medium

The invention relates to an interaction method and device based on virtual reality, equipment and a medium. According to the method, firstly, state data of a user in a virtual reality environment is collected and preprocessed to generate a standardized data stream, and then the distance, orientation similarity and interaction frequency between the user and a virtual avatar are calculated based on the data so as to construct a situational social graph representing social relation strength. Meanwhile, an attention model representing attention weight is constructed by analyzing a fixation point and a head direction of a virtual avatar of the user, and then information priority is calculated through weighted summation of social relation strength and the attention weight by utilizing a situational social graph and the attention model; according to the method and the device, the priority of the user is obtained, the audio, visual and text information is dynamically filtered or enhanced according to the priority to generate the optimized information flow, and finally the information flow is presented in the virtual reality environment, so that the cognitive load of the user in a dense social scene is effectively reduced, and the social interaction efficiency and immersion are improved.
Owner:SHIJIAZHUANG UNIVERSITY

Multi-dimension-based copywriting output method and device and related equipment

The invention discloses a multi-dimension-based copywriting output method and device and related equipment, and the method comprises the steps: obtaining the input information of a user, and judging the intention of the user based on the input information; if the user intention is a chat intention, generating a topic preference score based on the input information and the historical text information, generating a short-term popularity score based on the input information and the historical text information, generating a topic depth score based on the input information and the historical text information, and generating a marked memory node based on the long-term memory content; generating a comprehensive strategy score based on the topic preference score, the short-term popularity score, the topic depth score and the marked memory node; determining a target topic node based on the comprehensive strategy score; and outputting copywriting information corresponding to the input information based on the target topic node. By adopting the method, the natural language conforming to the context logic and the user preference can be output based on the copywriting input by the user, and the intelligent dialogue experience of the user is improved.
Owner:STAR CREATIVE ARTS (KUNSHAN) ENTERTAINMENT CO LTD

Digital employee auxiliary method and system for new energy station operation safety control

The invention discloses a new energy station operation safety control digital employee auxiliary method and system. The method comprises the following steps: automatically generating a workflow according to a new energy station field operation work ticket, an operation ticket and a measure list uploaded by a client, and issuing a corresponding new energy station inspection task to an associated client; constructing a multi-modal relation model based on the operation work ticket, realizing real-time sharing of a new energy station field operation video between clients by using live broadcast stream pushing, and identifying text information in the field operation video through OCR to confirm an operation position; and after the field operation of the new energy station is completed, performing new energy station operation result checking based on the relation model through image transmission, and outputting an alarm prompt when a checking error occurs. And state differences of the new energy station field equipment before and after operation are compared through image identification, state reset checking of the new energy station field operation equipment is completed, and a work report is generated and exported.
Owner:BEIJING SIFANG JIBAO AUTOMATION +1

Using Artificial Intelligence to Generate 3D Artifacts and Model Based Definition from 2D Drawings

Generating a 3D model from 2D drawings is provided. The method comprises extracting, by a design parser, content from 2D engineering drawings of an assembly and comparing the extracted content to a bill of materials corresponding to a 3D computer assisted design (CAD) model of the assembly to identify missing components from the 3D CAD model. Responsive to identifying missing components, 3D representations of the missing components are modeled based on the 2D engineering drawings, and metadata and textual information related to the 2D engineering drawings. The 3D representations of the missing components are incorporated into the 3D CAD model of the assembly to create a complete 3D CAD model. A manufacturing process for the assembly is then controlled according to the complete 3D CAD model.
Owner:THE BOEING CO

Unstructured document OCR error correction method and system based on multi-modal large model

The invention discloses an unstructured document OCR (Optical Character Recognition) error correction method based on a multi-modal large model, which comprises the following steps: step S101, acquiring an unstructured document to be processed, and decomposing the document into a page image unit and a preliminary optical character recognition text unit corresponding to the page image unit; step S102, selecting an error correction page, combining a page image unit with a preliminary optical character recognition text unit corresponding to the page image unit, and constructing a multi-modal input containing visual information and text information; step S103, providing the multi-modal input and a preset composite cue word to at least one multi-modal large model; step S104, the multi-modal large model generates a structured error correction suggestion according to the multi-modal input and the composite cue word; s105, displaying the original page image and the text attached with the error correction suggestions to the user on the interactive interface, receiving the processing operation of the user on each error correction suggestion, and recording the operation of the user as feedback data; and S106, optimizing an error correction system by using the feedback data.
Owner:SHANGHAI MITSUBISHI ELEVATOR CO LTD

Social media data-driven urban ponding depth inversion and precise positioning method

The invention provides a social media data-driven urban ponding depth inversion and precise positioning method. The method comprises the following steps of: 1, capturing and preprocessing multi-source social media data; step 2, performing ponding depth inversion based on image recognition; 3, accurately positioning the ponding position based on the text and the large model; extracting character information in pictures and videos by means of a picture character recognition technology in a large model, recognizing typical buildings therein, and assisting in positioning a flood position; screening out effective place information and time data related to ponding by adopting a large model; the identified location information is converted into specific location coordinates and detailed addresses by combining a geocoding technology, so that high-precision positioning is realized; and 4, outputting a comprehensive result: generating structured disaster information, and supporting direct calling of an emergency platform. Through combination of social media data and a space-time distribution model, ponding depth inversion is realized, quantitative data of the ponding depth is acquired, a disaster site is rapidly positioned, and decision making and flood control and disaster relief work are supported.
Owner:CHINA YANGTZE POWER +2

Intelligent conference summary generation method and device, equipment and storage medium

The invention provides an intelligent conference summary generation method and device, equipment and a storage medium, and relates to the technical field of natural language processing. According to the method provided by the invention, clear capture and accurate identification of audios are realized by using the multi-channel microphone array; and inputting the audio with the speaker identifier into the speech recognition model, and converting the audio into a structured initial summary text with a timestamp. Performing semantic analysis on the initial summary text through a large language model, and extracting key text information; a rule engine and a large language model are combined for combined judgment, so that the information accuracy is ensured; the initial summary text is converted into a first summary text in a preset text style, so that the text specialty is improved; humanized feature injection enhances the language naturalness of the second summary text; and the second summary text is converted into the structured target summary text, so that the conference summary better meets the requirements of professional scenes such as financial risk control conferences and medical consultation conferences, and the accuracy and availability of intelligent generation of the conference summary are improved.
Owner:PING AN TECH (SHENZHEN) CO LTD

Automatic game testing method and system

The invention provides an automatic game testing method and system, and belongs to the technical field of game development. The method comprises the following steps: acquiring image information and character information of a game; generating a prompt according to the acquired image information and text information; inputting the prompt into a preset model, and obtaining an operation instruction returned by the preset model; analyzing the operation instruction, and executing a corresponding game operation according to an analysis result; and collecting game image information and text information at the next moment after game operation is executed, and repeating the operation. According to the method, automatic testing of the game can be achieved, the flexibility of game testing is improved, and the cost of game testing is reduced.
Owner:GUANGZHOU AIYOU INFORMATION TECH

Character recognition management method and system fusing multi-modal data

The invention provides a character recognition management method and system fusing multi-modal data, and relates to the technical field of data processing, and the method comprises the steps: continuously capturing heterogeneous multi-source unsteady business instruction sequences submitted by employees; based on the heterogeneous multi-source unsteady service instruction sequence, executing a tensor reconstruction operation of an embedded space through a cross-modal depth feature alignment network, and outputting a structured high-dimensional tensor of a dimension specification; and based on the structured high-dimensional tensor, performing homogenized set division by using a K-means clustering algorithm driven by feature similarity to generate a feature set cluster with a compact feature. According to the invention, efficient management of text information in a banking business scene is realized.
Owner:JIANGXI BANK CO LTD

screen module

1. Name of the designed product: screen module. 2. Use of the designed product: for displaying image, text information and voice information processing. 3. Design points of the designed product: in shape. 4. Picture or photo best indicating the design points: perspective view.
Owner:QINDAO HAIER REFRIGERATOR CO LTD +1

Game machine

To provide a game machine capable of increasing working efficiency in a manufacturing process of the game machine by a worker having a visual handicap.SOLUTION: In a Pachinko machine equipped with a game board having a flow-down area in which game balls can flow down, a recessed and projecting notation section 2111 indicated in a recessed and projecting shape is provided in a predetermined position of a middle flow path section 2102 provided in the game board as a flow path of game balls, character information by the recessed and projecting notation section 2111 is provided such that it can be touched by finger tips, the middle flow path section 2102 is mounted to a ball route unit 2100A provided in a predetermined position within the game board, and the recessed and projecting notation section 2111 cannot be visually recognized easily from the outside than when it is visually recognized directly when it is correctly mounted to the ball route unit 2100A, and the recessed and projecting notation section 2111 is visually recognized easily from the outside when it is mounted wrongly than when it is correctly mounted to the ball route unit 2100A.SELECTED DRAWING: Figure 160
Owner:DAIICHI SHOKAI KK

Electric power communication character extraction method and device, storage medium and computer equipment

The invention discloses an electric power communication character extraction method and device, a storage medium and computer equipment, relates to the technical field of character information extraction, and mainly aims to improve the extraction efficiency and extraction precision of electric power communication characters. The method comprises the following steps: acquiring a to-be-identified electric power communication text image; obtaining a preset character extraction model, and performing compression processing on the preset character extraction model to obtain the compressed preset character extraction model; and inputting the to-be-identified electric power communication text image into the compressed preset character extraction model for character extraction to obtain characters corresponding to the to-be-identified electric power communication text image. The method is suitable for the scene of extracting the electric power communication characters.
Owner:EAST CHINA BRANCH OF STATE GRID CORP

Efficient semantic perception pre-training track representation learning method and device

The invention discloses an efficient semantic perception pre-training track representation learning method and device, and relates to the technical field of artificial intelligence, and the method comprises the steps: generating track embedding of a vehicle track through a first encoder; training a first encoder according to the road view and the interest point view of the vehicle trajectory; initializing a second encoder according to the weight of the first encoder; generating a compressed embedding of the vehicle trajectory by a second encoder; and aligning the track embedding with the compression embedding. Through travel purpose perception pre-training, a large language model does not need to be introduced in a downstream task coding stage, the travel purpose is effectively captured, additional calculation overhead is avoided, the problem that calculation burden is heavy when text information is integrated through a traditional method is solved, semantic information in a vehicle track can be efficiently and fully learned, and the method is suitable for being used for a vehicle. And possibility is provided for real-time trajectory analysis.
Owner:BEIJING JIAOTONG UNIV