Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

689 results about "Optical character recognition" patented technology

Optical character recognition or optical character reader (OCR) is the electronic or mechanical conversion of images of typed, handwritten or printed text into machine-encoded text, whether from a scanned document, a photo of a document, a scene-photo (for example the text on signs and billboards in a landscape photo) or from subtitle text superimposed on an image (for example from a television broadcast).

Engineering document index consistency proofreading method and system based on multi-modal large model

The invention relates to an engineering document index consistency proofreading method and system based on a multi-modal large model, and the method comprises the steps: Q1. OCR detection and recognition: carrying out the optical character recognition and format analysis of a source document, converting an uploaded PDF document into a processable text message in a Markdown format, and carrying out the format discrimination of a table, a formula and a plain text; and Q2, table and formula processing: adopting a hierarchical processing strategy, intelligently selecting an optimal processing mode according to the complexity of the table, and converting table information into a descriptive long text through a language large model and cue words. According to the method, accurate, reliable and efficient document index checking service can be provided for a user, the quality and efficiency of professional document processing are remarkably improved, the efficiency and quality of knowledge graph construction are remarkably improved, a knowledge verification system capable of being evolved continuously is established, and the method is suitable for popularization and application. And a reliable technical support is provided for knowledge management and professional decision-making in a complex field.
Owner:CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719

RAG-based voucher classification method, medium and equipment

The invention relates to an RAG-based voucher classification method, a medium and equipment, and the method comprises the steps: receiving digital image data of a to-be-classified voucher, carrying out the multi-modal optical character recognition processing to generate a structured OCR result, extracting a text semantic feature vector and a visual layout feature vector based on the structured OCR result, carrying out the fusion of the text semantic feature vector and the visual layout feature vector to generate a multi-modal query vector, and carrying out the classification of the to-be-classified voucher. Similar samples and semantic similarity scores and category metadata thereof are obtained through approximate nearest neighbor retrieval, after an initial candidate category list is generated, key field values are extracted for each candidate category, evidence credibility scores are calculated, comprehensive confidence scores are generated by fusing the semantic similarity scores and the evidence credibility scores, reordering is conducted, and a candidate category list is obtained. And finally, selecting a classification decision path according to the score distribution, and outputting a classification result and an interpretability report. The accuracy and robustness of voucher classification are effectively improved, and complex voucher scenes with changeable formats and fuzzy semantics can be processed; and the interpretability and reliability of the classification decision are enhanced.
Owner:FUJIAN BOSS SOFTWARE

Multi-mode intelligent auditing method for supply chain purchase bid invitation project

The invention provides a supply chain purchase bid invitation project multi-mode intelligent auditing method, and belongs to the technical field of electronic purchase. Comprising the following steps: defining a review rule system comprising rule contents, review points and review logic; in the bidding stage, bidding files and quotation information uploaded by suppliers are received, and after bid opening, bid invitation files, bidding files of entered suppliers and corresponding review rule systems are input into a multi-mode large model server; performing unstructured data analysis on the bidding document and the bidding document by using a multi-modal large model in combination with an OCR (Optical Character Recognition) technology, and identifying and matching response contents in the bidding document through intention; pushing the extracted content to a corresponding business auditing agent, and performing judgment according to an auditing point and auditing logic; and summarizing the review results, and generating an intelligent review report containing the number of conformity items, the number of non-conformity items, conclusions of the review points and judgment bases. The examination efficiency and accuracy are improved, and human errors and compliance risks are reduced.
Owner:INSPUR GENERSOFT CO LTD

Device for time-based tracking and cost optimization in construction projects

A device for time-based tracking and cost optimization in construction projects, the device comprising the following: a robust housing suitable for use on construction sites; a processing unit located inside the housing, configured to perform real-time time-stamping, data acquisition and preprocessing tasks; a multimodal sensor unit that is operationally coupled with the processing unit, wherein the sensor unit comprises at least a motion sensor, an RFID reader, sensors for environmental conditions and a vision module with optical character recognition; a real-time clock module that is operationally connected to the processing unit to provide time synchronization for all sensor data streams; a wireless communication module that supports the Wi-Fi, LoRa and LTE protocols and is configured for transmitting time-stamped data to a central project server; a storage module that is operationally coupled with the processing unit to locally buffer time series data of construction activity during offline operation; a housing-mounted, touchscreen-based human-machine interface configured to allow site personnel to enter activity updates and confirm the status of construction tasks; a cost optimization engine running on the central server, the engine being configured to receive time-synchronized sensor data from multiple such devices and dynamically calculate time-cost trade-offs using a predictive planning technique that incorporates the principles of the critical path and the power value; furthermore, the device is configured to be integrated into a digital twin environment of the building under construction in order to provide real-time visualization of progress and to generate suggestions for resource reallocation based on a time-cost-benefit analysis.
Owner:1XL INFRA & REAL ESTATE DEVELOPMENT LLC +2

Document table extraction method and device, equipment and medium

The invention discloses a document table extraction method and device, equipment and a medium, and relates to the technical field of computer information processing. The extraction method comprises the following steps: performing OCR (Optical Character Recognition) on a to-be-processed document table image to obtain a text block; performing visual feature coding on the document table image to obtain deep visual features; performing semantic feature coding on the text sequence of the text block to obtain a semantic feature vector; performing spatial feature coding on the bounding box of the text block to obtain a spatial feature vector; performing feature fusion processing on the deep visual features, the semantic feature vectors and the spatial feature vectors to obtain multi-modal guide features; and performing structured decoding processing on the multi-modal guide features to obtain structured representation of the table. According to the method, the text and the position information pre-recognized by the OCR are fused with the visual features of the document table, so that the visual features are guided to be expressed again and are actively aligned to the logic structure defined by the prior information, and the extraction accuracy of the table logic structure is improved.
Owner:SICHUAN ENRISING INFORMATION TECH CO LTD

Nuclear power safety report data extraction method and system based on multi-modal feature fusion

The invention provides a nuclear power safety report data extraction method and system based on multi-modal feature fusion, and belongs to the technical field of data processing, and the method comprises the steps: obtaining a nuclear power safety report document; carrying out OCR (Optical Character Recognition) and paragraph segmentation on the report document to obtain first text data; performing text semantic understanding on the first text data by adopting a dynamic template matching and semantic driving extraction mechanism to obtain a text feature vector; segmenting pages of the report document and identifying fonts and fonts to generate first format data; adopting a cross-page table reconstruction algorithm to detect table cells of the first format data and splicing a cross-page table to obtain format feature vectors; performing feature alignment on the text feature vector and the format feature vector, and inputting the text feature vector and the format feature vector into a large pre-training language model to obtain a semantic vector; and fusing the text feature vector, the format feature vector and the semantic vector to form a composite expression unit, and intelligently extracting structured data based on the composite expression unit. The semantic entity and structural relationship in the nuclear power report can be effectively identified.
Owner:SHANGHAI NUCLEAR ENGINEERING RESEARCH & DESIGN INSTITUTE CO LTD

System and methods for document processing for data extraction and matching

System and methods are disclosed for matching extracted text data based on one or more similarity scores. The method may include receiving one or more documents from a plurality of data sources, utilizing an optical character recognition algorithm for extracting text data from the one or more documents, comparing, utilizing a fuzzy matching algorithm, the extracted text data to reference dataset(s) to determine one or more matches between the extracted text data and at least one of the reference dataset(s), wherein the one or more matches are based on at least one similarity score, inputting the determined one or more matches and the at least one similarity score into a trained machine-learning model to refine the one or more matches, and outputting a representation of the refined one or more matches and the at least one similarity score to a graphical user interface of a device.
Owner:STATE FARM MUTAL AUTOMOBILE INSURANCE COMPANY

Document interpretation and report generation method and device, equipment and medium

The invention relates to the technical field of natural language processing, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a document interpretation and report generation method, device, equipment and medium, which comprises the following steps: receiving an original document set to generate a structured document object, executing optical character recognition on an image content set to generate a recognition text set, the recognition text set and the text content set are combined into a unified text sequence, element item extraction is executed based on the interpretation template parameter set to generate an interpretation element set, a retrieval enhancement context is retrieved and generated from the domain knowledge base, and the unified text sequence, the interpretation template parameter set and the retrieval enhancement context are input into a language model to generate an interpretation result. And generating report content based on the historical report template set. According to the method, automatic closed loop of document interpretation and report generation is realized through multi-modal unified processing and semantic enhanced reasoning, the efficiency is improved, and the manual dependence and compliance risk are reduced.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Multi-modal information analysis and scheme reminding method and device, equipment and medium

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a multi-modal information analysis and scheme reminding method, device, equipment and medium, which comprises the steps of receiving input data and converting the input data into multi-modal data, executing optical character recognition and image classification recognition to generate a recognition result, and sending the recognition result to a server; a natural language processing model is used for analyzing fuzzy description to generate an analysis result, a knowledge base is inquired, a knowledge graph is combined to generate an association result, the analysis result, the association result and user feature data are fused to generate an execution scheme, the execution scheme is compared with an abnormal list, supervision confirmation is triggered, and a compliance instruction is generated. And personalized reminding contents are generated. The information analysis integrity is improved through multi-modal recognition, natural language processing and the knowledge graph, supervision confirmation and personalized reminding are introduced, and intelligent, compliant and reliable reminding management is achieved.
Owner:PING AN TECH (SHENZHEN) CO LTD

OCR (optical character recognition) method and system for high-precision table data structuring

The invention provides an OCR (Optical Character Recognition) method and system for high-precision table data structuring. The method comprises the following steps: converting an original image into a grayscale image and preprocessing the grayscale image; extracting a table edge structure line and filling a fracture part; detecting longitudinal and transverse straight lines in the preprocessed image, calculating intersection points, and determining a table row-column structure; dividing a cell region and positioning to generate a cell coordinate matrix; pixels in the cells are divided into a frame influence area and an effective data area, and the frame influence area executes neighborhood mean filtering and weighted fusion operation; performing end-to-end detection on characters and symbols in the effective data area, and outputting an OCR recognition result with coordinates; and dynamically generating a table structure template based on the cell coordinate matrix and an OCR recognition result, speculating a strategy matching field type through a rule, processing and merging cell missing data based on adjacent cell information, and outputting structured data. The reliability and accuracy of the OCR technology are improved, and the requirement of automatic information processing for high-precision data extraction is met.
Owner:WUHAN UNIV

Industrial drawing analysis method and system combining multi-modal large model and OCR (optical character recognition)

The invention discloses an industrial drawing analysis method and system combined with a multi-modal large model and OCR, and relates to the technical field of drawing analysis, the method comprises the following steps: carrying out layout area segmentation, text recognition and geometric element extraction on an industrial drawing to obtain a structured information set; constructing a multi-relation structure chart set of the industrial drawings; performing structure embedding and feature bias enhancement to obtain a structure feature set; performing semantic understanding on the industrial drawing to obtain semantic features, and performing multi-channel coding to obtain a multi-modal feature set; obtaining a preliminary analysis result set of the industrial drawing; and performing rule verification, and outputting an industrial drawing analysis result. The technical problems of low industrial drawing analysis efficiency, inaccurate automatic scheme analysis and limited processing capacity in the prior art are solved, and the technical effects of realizing full-process automatic analysis of the industrial drawing by combining the multi-modal large model and the OCR, improving the precision and efficiency of industrial drawing analysis and standardizing output information are achieved.
Owner:SUZHOU DEMI TECHNOLOGY CO LTD

PDF document intelligent word segmentation method based on content structure perception

The invention provides a PDF document intelligent word segmentation method based on content structure perception, and belongs to the technical field of computer information.The PDF document intelligent word segmentation method comprises the steps that firstly, by fusing layout analysis and multi-tool analysis, multiple types of content such as title levels, ordinary tables, characters in pictures and tables in the pictures in a PDF are accurately recognized and extracted; secondly, providing a structure-perceived word segmentation strategy, and carrying out differentiation processing according to content types, namely, dividing semantic blocks by taking a title as a guide, taking a complete table as an independent semantic unit, and carrying out context association on a picture OCR (Optical Character Recognition) result; and finally outputting high-quality knowledge fragments rich in metadata such as levels and types. According to the method, the information integrity and semantic accuracy of the PDF document during construction of the large model knowledge base can be remarkably improved.
Owner:浪潮智慧城市科技有限公司

Test paper correction and personalized learning method based on OCR (Optical Character Recognition) and large language model

The invention provides a test paper correction and personalized learning method based on OCR and a large language model. The method comprises the steps of S1, test paper scanning and preprocessing; s2, performing OCR (optical character recognition) and data structuring; s3, intelligently correcting the large model; s4, performing multi-dimensional statistical analysis; and S5, generating and pushing variable questions. According to the method, technologies such as OCR (Optical Character Recognition), a large language model and a knowledge graph are fused, so that the whole process intelligence from test paper correction to personalized learning guidance is realized, the efficiency and accuracy of education evaluation are improved, and technical support is provided for personalized teaching.
Owner:AZURE ORIGIN SMART TECHNOLOGY (HANGZHOU) CO LTD

Dynamic double-layer hidden watermark and encryption binding file protection method and system based on deep learning

The invention relates to a dynamic double-layer hidden watermark and encryption binding file protection method and system based on deep learning, and belongs to the technical field of digital content security. The problems of attack resistance, traceability obstruction and key-watermark unhooking in document cross-platform circulation are solved. According to the scheme, the method comprises the following steps of: extracting semantic fingerprints by using a sentence vector model Sentence-BERT; the fuzzy extractor generates a master key and derives a time key chain; the authentication encryption algorithm AEAD encrypts and binds the source and the timestamp load; container layer structure rearrangement and document layer zero-width character double embedding are carried out; the generative adversarial network or diffusion model adversarial training improves the optical character recognition and transcoding resistance; version binding and tracing are achieved through the watermark hash chain. The technical effects cover anti-counterfeiting migration, cross-layer fault-tolerant guarantee recoverability, rearrangement attack resistance, full-period accurate traceability and post-quantum security enhancement.
Owner:SOUTHWEST UNIV

Multi-dimensional automatic background investigation method and system based on Ai assistance

The invention relates to the technical field of human resources, and discloses a multi-dimensional automatic background investigation method and system based on Ai assistance. The system comprises a multi-source data acquisition module, an intelligent analysis engine module, a real-time interview processing module, a quality verification control module, a dynamic decision support module, a privacy security protection module and a visual report generation module. According to the method, public data of candidates are acquired through a multi-thread crawler technology, and the authenticity of educational background is verified; analyzing the paper certificate by adopting an optical character recognition technology; constructing a space-time relation knowledge graph to detect work experience abnormality; interview content is transcribed in real time, and emotional tendency is analyzed; dynamically adjusting the evaluation dimension weight; and generating a visual risk assessment report. According to the method, the enterprise employment risk is reduced, the comprehensiveness and depth of background investigation are improved, the accuracy and applicability of a background investigation conclusion are improved, a safe and reliable background investigation environment is constructed, and the enterprise human resource allocation effect is optimized.
Owner:CAIXING (GUANGZHOU) TECH SERVICE CO LTD

Identifying Items in Images Using Embeddings Generated from the Images and Ranking Candidates Using a Language Model

An online system applies a visual language model and an optical character recognition model to a received image to generate descriptive information about unknown items in the image. The online system prompts a generative model with the descriptive information about unknown items in the image to separate the descriptive information into different bins each corresponding to a different unknown item in the image. For each unknown item detected in the image, the online system generates a target embedding from its descriptive information and performs a nearest neighbor search on an item catalog including embeddings for various items to find a set of candidate embeddings matching the target embedding. The online system retrieves item attributes of candidate items each corresponding to a candidate embedding of the set and prompts the generative model with this information to rank candidate items for the unknown item in the image.
Owner:MAPLEBEAR INC

PDF (Portable Document Format) document structured extraction system based on multi-modal language model

The invention discloses a PDF (Portable Document Format) document structured extraction system based on a multi-modal language model, belongs to the technical field of document processing and optical character recognition, and aims at solving the technical problem of how to improve the existing OCR (Optical Character Recognition) technology to improve the analysis capability of a complex document structure and improve the recognition precision of handwritten forms and other non-standard fonts. According to the technical scheme, the system adopts a layered decoupling architecture and comprises an input layer, a preprocessing layer, a reasoning layer, an output layer and a monitoring and fault-tolerant module; wherein the output layer is used for multi-source data access and path management to realize a local file system or S3 cloud storage; the preprocessing layer is used for invalid document filtering and visual feature extraction; the reasoning layer is used for multi-modal model interaction and content processing; the output layer is used for outputting a content aggregation result; and the monitoring and fault-tolerant module is used for realizing real-time state monitoring, resource consumption analysis and exception handling.
Owner:JIANGSU HAIRUO INFORMATION TECHNOLOGY CO LTD

Automatic software testing method and system based on multi-mode AI cooperation

The invention relates to the technical field of automatic testing, in particular to an automatic software testing method and system based on multi-mode AI cooperation. According to the scheme, optical character recognition, voice recognition and natural language processing analysis tools are packaged into independent containers, and analysis results are input into a multi-mode semantic association model in real time; calculating the semantic similarity between the image description text and the corresponding voice transcription text; constructing a test demand complexity evaluation model; defining use case quality indexes, and establishing an error sample library to store manually labeled problem use cases and correction labels thereof; the intention demand is finally determined according to the screening result of the three-stage filter; and setting a test case template, optimizing a path according to the fusion quality, improving the case generation efficiency, and synchronously updating the optimized path to a knowledge base associated with the error sample library. According to the scheme, automatic software testing is achieved through multi-modal AI cooperation in combination with containerized deployment, model dynamic optimization and the like, and efficiency and recognition accuracy are improved.
Owner:RUIJIAN TECHNOLOGY (BEIJING) CO LTD

Information extraction from unstructured documents with hybrid retrieval augmentation using multi-modal language models

A system for extracting a number of data elements from one or more data sources. Image-based documents are indexed using optical character recognition and a text embedding model to convert the document text to vector embeddings. Relevant portions of the document are identified by comparing the vector embedding of the documents to a vector embedding of a prompt or a request to extract information. The relevant text is mapped to a corresponding page of the documents. The page may be provided to a multi-modal language model for information extraction. The multi-modal language model can process contextual information included in the layout, figures, markings, etc. of the document to extract the information. The system populates an ontological data store based on the response from the language model. Extraction accuracy is improved without significant increases in computations performed by the system.
Owner:AMERICAN INTERNATIONAL GROUP INC

UI automatic test method and test system, and storage medium

The invention provides a UI automatic test method and test system and a storage medium. The method comprises the steps that an optical character recognition result of a target text and a to-be-tested UI is obtained, and the optical character recognition result comprises a plurality of elements and element texts, coordinate frames and visual features of all the elements; obtaining the semantic similarity between the target text and each element text; obtaining attribute similarity between the target text and each element based on the visual features of each element; according to the semantic similarity and the attribute similarity, elements, corresponding to the target text, in the optical character recognition result are determined to serve as target elements; taking the position of the coordinate frame of the target element in the UI as a target position; and generating a presentation image which highlights the target position and / or the target element. Therefore, fuzzy matching based on semantics can be realized, multi-modal features including text semantics and visual features are fused, and the robustness and accuracy of UI element positioning are improved in combination with context sensing.
Owner:AIJI MICRO CONSULTING (XIAMEN) CO LTD

Multi-modal document data processing method and system oriented to large language model training

ActiveCN121093293ANeural learning methodsBatch processingCharacter (computing)
The invention discloses a multi-modal document data processing method and system for large language model training, and the method comprises the steps: receiving a plurality of original documents in various formats, extracting the structure information of each original document, and recognizing a text region and an image region of each original document based on the structure information; performing optical character recognition on the text region and the image region by adopting a parallel OCR (Optical Character Recognition) engine based on GPU (Graphics Processing Unit) acceleration and heterogeneous calculation to generate recognition text data of the corresponding original document; performing multi-dimensional quality evaluation and cleaning on the recognition text data of each original document, and outputting normalized text data; and storing the standardized text data into a distributed knowledge base according to a predefined structure, and performing copyright and compliance test on the standardized text data. By adopting a parallel OCR recognition engine based on GPU acceleration and heterogeneous calculation, efficient and high-precision batch processing of multi-modal documents is realized, and the processing speed, the recognition accuracy and the data quality are improved.
Owner:HANGZHOU BINGTE TECH

Document structure extraction and model training method and device, equipment and medium

The invention discloses a document structure extraction and model training method and device, equipment and a medium, and relates to the technical field of artificial intelligence and computer vision. The method comprises the following steps: constructing a special training data set containing data of at least two document understanding tasks (including optical character recognition, layout analysis, text positioning, regional text extraction, image description and chart title generation), and a fine tuning data set for converting a document image into a machine-readable structured text format; constructing a multi-modal large model comprising a shape adaptive cutting module, a visual encoder, a visual token compression module, a modal connector and a language decoder; pre-training the model by using the special training data set to jointly learn various document understanding tasks; and performing fine tuning on the pre-trained model by using the fine tuning data set, and adapting to a document structure extraction task to obtain a document structure extraction model. By means of the technical scheme, efficient and accurate document structure extraction can be achieved.
Owner:CETC CYBERSPACE SECURITY TECH CO LTD

Video character recognition and erasing method and system, storage medium and electronic device

The invention discloses a video character recognition and erasing method and system, a storage medium and an electronic device. The method comprises the following steps: acquiring a video; performing frame extraction processing to obtain a video frame picture; the method comprises the following steps: acquiring all text contents and coordinate data in a picture through OCR (Optical Character Recognition) identification, analyzing and identifying flower characters and subtitles on a video frame picture through a multi-modal large model, screening out a region of the flower characters and the subtitles needing to be erased through coordinate matching, and determining the region as an erased region; and intelligently erasing and repairing the flower and subtitle areas by adopting a video repairing technology, recovering the original state of the video, and generating an erased new video file. According to the method and the device, the full process of automation is realized, subtitles and flower characters in the video do not need to be manually marked, identified and erased, printed characters of commodities / packages are reserved, mistaken erasure is avoided, and the processing efficiency of video reediting is greatly improved.
Owner:GUANGZHOU KUAIZI INFORMATION TECH CO LTD

Multi-mode intelligent semantic understanding and abstract generation system and method based on HDMI (High Definition Multimedia Interface) stream

The invention provides a multi-mode intelligent semantic comprehension and abstract generation system and method based on an HDMI stream, and the system comprises an HDMI input module which is used for receiving a video signal outputted by external equipment; the image content streaming analysis module is configured to perform content segmentation, optical character recognition, layout structure extraction and graphic element detection on the video signals and output structured image data; the audio acquisition module is used for acquiring an audio signal and preprocessing the audio signal; the voice recognition module is used for transferring the preprocessed audio signal into a voice text; and the multi-modal fusion and semantic understanding module is configured to fuse the structured image data and the voice text, generate summary information and output the summary information to the result output module. According to the invention, real-time synchronous analysis of the multi-modal content and intelligent generation of the abstract are realized.
Owner:BEIJING YUNJIANXIN TECH CO LTD

Medical document intelligent identification method and system based on OCR (Optical Character Recognition)

The invention discloses an OCR-based medical document intelligent identification method and system, and relates to the technical field of document identification, and the method comprises the steps: collecting a medical document image through a mobile terminal, generating a binary image based on the collected image through U-Net in combination with multi-scale feature fusion and an attention mechanism, and cutting the binary image; based on the cut binary image, text information is extracted through an OCR model, and a structured field is extracted according to the typesetting rule and geometric distribution of the medical document; and performing deterministic rule judgment and risk assessment on the structured data. According to the method, the image calculation complexity is reduced through standard graying processing, the text region feature extraction precision is improved by fusing a U-Net structure of a CBAM attention mechanism, effective fusion and noise suppression of multi-scale features are realized in combination with Attention Gate, text direction correction is realized in combination with Hough transform, and the text detection robustness and recognition accuracy are improved.
Owner:SHALLBRIGHT HEALTHTECH CO LTD

Knowledge base context awareness and traceability enhanced intelligent retrieval and question-answering system

The invention provides an intelligent retrieval and question-answering system for knowledge base context awareness and traceability enhancement, and the system obtains an analysis result, an optical character recognition result and a semantic analysis result through the analysis of PDF physical layout, Word / Excel paragraphs, titles, tables and style attributes thereof, and the optical character recognition. Structured knowledge blocks, positions and levels of path images and vector representation are generated and stored in a vector database, and related results of user query requests are extracted by executing a mixed retrieval algorithm with keyword retrieval and vector semantic retrieval. Performing intelligent reordering by considering semantic similarity, keyword matching, knowledge block type weight, source knowledge base weight and page position weight to obtain a candidate knowledge block list, constructing cue words according to an intelligent reordering result, and calling an external large language model to generate answers; source labels in answers are managed to be associated with metadata of corresponding numbers in a candidate knowledge block list, and the problem that text blocks and traceability are not accurate is solved.
Owner:CHINA HAISUM ENG

System And Method for Generating or Augmenting Lighting Effects Based on Contextual Data

The present invention relates to a system and method for generating lighting effects based on contextual data. The system includes one or more sources, including microphones, cameras, biometric sensors, environmental sensors, and digital feeds providing information including time, date, season, geolocation, weather, calendar events, or current events. A processor analyzes the contextual data to determine a contextual state associated with an ongoing event or user condition. The processor further identifies secondary trigger conditions through computer vision analysis of screen content, developer-scripted events, generative AI-derived cues, optical character recognition, gameplay events, or transformation of on-screen 2D or 3D visual regions. Based on the contextual or composite contextual state, the processor determines one or more lighting effects, which are then triggered on one or more output devices. This enables dynamic and contextually adaptive lighting responses to both real-world and digital stimuli.
Owner:WHIRLWIND VR INC

Classroom teaching evaluation generation method and system based on multi-modal large model

The invention provides a classroom teaching evaluation generation method and system based on a multi-modal large model, and belongs to the technical field of artificial intelligence and education, and the method comprises the steps: S1, defining an analysis task; s2, setting unified prompt words for guiding model generation: role setting, core tasks and structure requirements; s3, performing fine adjustment on the detector by using the corresponding task domain target detection data set; and S4, integrating a detector, a region slice detection strategy, an optical character recognition technology, a voice recognition technology and a framework of a plurality of large models, taking a teaching video and a corresponding audio as input of the classroom teaching evaluation model, and outputting a structured teaching evaluation analysis report. According to the method provided by the invention, a better effect is obtained on various evaluation standards, the quality of an evaluation report is higher, and the method can be expanded to any classroom scene only by replacing or finely adjusting a detector capable of adapting to a new scene, so that efficient and low-cost framework migration is realized, and meanwhile, high performance is ensured.
Owner:BEIHANG UNIV

Training and using a vector encoder to determine vectors for sub-images of text in an image subject to optical character recognition

Provided are a computer program product, system, and method for training and using a vector encoder to determine vectors for sub-images of text in an image to subject to optical character recognition. A vector encoder is trained to encode images representing text into vectors in a vector space. Vectors of images representing similar text have a high degree of cohesion in the vector space. Vectors of images representing dissimilar text have a low degree of cohesion in the vector space. An input image is processed to determine sub-images of the input image that bound text represented in the input image. The sub-images are inputted to the vector encoder to output sub-image vectors. The vector encoder generates a search vector for search text. Optical character recognition is applied to at least one region of the input image including the sub-images having sub-image vectors matching the search vector.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Image processing-based aeronautical part identification character recognition method and system

The invention relates to the technical field of image processing, in particular to an aeronautical part identification character recognition method and system based on image processing, and the method comprises the steps: obtaining a character image of an aeronautical part, carrying out the stroke fracture reconstruction of a stroke fracture in the character image, so as to generate a reconstructed image, counting the number of connection pixel points newly added in the stroke fracture reconstruction process; performing optical character recognition on the reconstructed image to obtain a preliminary recognition result and a corresponding matching confidence coefficient; determining a stroke integrity index of the character image based on the number of the newly added connection pixel points and the total number of pixel points of the reconstructed image; the matching confidence coefficient of optical character recognition and the stroke integrity index based on reconstruction pixel quantization are combined, double judgment is carried out, it is ensured that the recognition result has the two characteristics of similar forms and reliable sources, and the accuracy and reliability of character recognition under the complex, high-light-reflection and abrasion conditions are improved.
Owner:HANZHONG QUNFENG MACHINERY MFG