Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

40 results about "Text annotation" patented technology

Text Annotation is the practice and the result of adding a note or gloss to a text, which may include highlights or underlining, comments, footnotes, tags, and links. Text annotations can include notes written for a reader's private purposes, as well as shared annotations written for the purposes of collaborative writing and editing, commentary, or social reading and sharing. In some fields, text annotation is comparable to metadata insofar as it is added post hoc and provides information about a text without fundamentally altering that original text. Text annotations are sometimes referred to as marginalia, though some reserve this term specifically for hand-written notes made in the margins of books or manuscripts. Annotations are extremely useful and help to develop knowledge of English literature.

Lead mark arrangement method and device, electronic equipment and storage medium

PendingCN121389217AGeometric CADConfiguration CADSoftware engineeringText annotation
The invention provides a lead mark arrangement method and device, electronic equipment and a storage medium, relates to the technical field of computers, in particular to the fields of engineering drawing, computer aided design and the like, and can be used for application scenes such as drawing lead mark arrangement and the like. According to the specific implementation scheme, the method comprises the steps of obtaining a plurality of to-be-arranged lead marks and a plurality of mark objects; aiming at each to-be-arranged lead mark, determining a lead-out point coordinate of a lead based on a corresponding mark object; in response to a text arrangement mode selected by a user, determining text coordinates according to the extraction point coordinates; and generating a lead mark arrangement result according to the lead-out point coordinate and the text coordinate corresponding to each to-be-arranged lead mark. According to the scheme, the lead text annotation arrangement can be quickly carried out, various arrangement effects can be obtained by setting a simple arrangement mode, and the operation efficiency of arrangement and annotation when a user makes a picture is greatly improved.
Owner:HANGZHOU QUNHE INFORMATION TECHNOLOGIES CO LTD

Speech recognition model training method and device, equipment and readable storage medium

PendingCN121662029ASpeech recognitionData packText annotation
The invention discloses a speech recognition model training method and device, equipment and a readable storage medium, and relates to the technical field of artificial intelligence. Comprising the following steps: firstly, acquiring voice training data and annotation data corresponding to the voice training data; the annotation data comprises text annotation data and intention annotation data; fuzzy processing is carried out on the text labeling data, and voice features of the voice training data are extracted; and training a speech recognition model based on the speech features, the text annotation data and the intention annotation data until the speech recognition model converges. According to the method, the text annotation data is fuzzified, some characters which are not concerned about intention classification are ignored, the speech recognition model is more focused on keywords, intention classification information is introduced in the training process, and the accuracy of the recognition result generated by the speech recognition model for intention classification is improved.
Owner:AISPEECH CO LTD

A method and system for recommending low-forgetting English annotation styles based on AHP.

ActiveCN120493881BReduce the rate of memory forgettingAuxiliary memoryData processing applicationsNatural language data processingLexicologyText annotation
This invention proposes a method and system for recommending English annotation methods with low forgetting rate based on the Analytic Hierarchy Process (AHP), belonging to the fields of optimization and intelligent decision theory. First, from the perspective of English learners, this invention analyzes and identifies five key factors affecting the rate of forgetting English vocabulary. Then, using the Analytic Hierarchy Process (AHP), a weight matrix that passes consistency detection is constructed for each influencing factor, and the actual influencing factor values ​​are normalized based on fuzzy logic. Based on the weight matrix that passes consistency detection and the normalized influencing factor scores under different text annotation methods, the estimated forgetting rate of English vocabulary for English learners under the corresponding annotation methods is obtained. The annotation method with the lowest estimated forgetting rate is selected as the optimal recommended annotation method to assist English learners in vocabulary learning. Finally, a one-month experimental verification of the proposed recommendation algorithm is conducted, and the experimental results demonstrate the effectiveness of this invention.
Owner:XIAN KEDAGAOXIN UNIV

A multi-screenshot editing and annotation system and a method of using the same

The application provides a multi-screenshot editing and annotation system and a use method thereof, which comprises: a screenshot module, which realizes execution area screenshot, full-screen screenshot or window screenshot operation; an interface control module, which realizes dynamic adjustment of main interface size and menu layout according to screenshot size, controls interface top edge display and shelter area avoidance; a puzzle and overlay module, which realizes zooming, rotating, mirroring or transparency transformation of screenshots and arranges them in an automatic puzzle or overlay mode; a graphic editing module, which realizes addition / modification of graphic elements, realizes picture comparison and CAD auxiliary design function; a graphic and text annotation module, which realizes addition of graphic symbol marks, text annotation and mechanical drawing standard annotation; and a fast picture file management module, which realizes automatic saving of screenshots according to serial numbers, quick browsing and modification and user self-defined setting management. The application provides a more efficient and convenient multi-screenshot editing and annotation system and method, so as to simplify operation process, improve work efficiency and enhance user experience.
Owner:SHENZHEN LIANYING TECH CO LTD

Nursing teaching task-oriented nursing field text annotation corpus construction method

The invention relates to the field of artificial intelligence technology and medical information processing, and discloses a nursing teaching task-oriented nursing field text annotation corpus construction method, which comprises the following steps of S1, collecting original nursing text data and constructing a training data set; s2, nursing text data cleaning and entity labeling standardization; s3, constructing an entity recognition model oriented to the nursing field; s4, calculating a total loss function based on the main loss and the auxiliary loss; s5, training the entity recognition model by adopting the cleaned and standardized training data set; and S6, carrying out automatic labeling on the original nursing text by utilizing the entity recognition model. The method has the beneficial effects that the nursing field dynamic dictionary is constructed, and the fuzzy matching function fusing the editing distance similarity and the semantic similarity is introduced, so that spelling errors, term variants and the like in the original text can be intelligently mapped to the standard words, deep standardized cleaning of the nursing text is realized, and the user experience is improved. And the problem of data noise is effectively solved.
Owner:TIANJIN TELLYES SCI INC +1

Bridge management and maintenance knowledge graph construction method, system and equipment based on large-model multi-agent and medium

The invention discloses a large-model multi-agent-based bridge management and maintenance knowledge graph construction method, system and device and a medium, and the method comprises the steps: processing a bridge inspection report, converting an original text into a unit through a decomposition agent, and building a text annotation database; the method comprises the following steps: taking a subject text block as input, extracting a triple by an extraction agent, then checking whether the extracted triple conforms to a bridge detection domain ontology and domain knowledge or not by a verification agent, finally correcting the triple with errors by a correction agent, iteratively extracting, verifying and correcting the triple, and outputting a high-quality triple; establishing a dynamic knowledge learning mechanism, and updating a knowledge base; constructing an initial knowledge graph according to the verified knowledge triad; and the review agent reviews the initial knowledge graph and feeds back a review result to the construction agent for iterative updating, and finally, the bridge management and maintenance knowledge graph is constructed. According to the method, the accuracy and the fact consistency of the atlas data can be ensured.
Owner:SOUTHEAST UNIV

Machine-readable constructional engineering standard digital processing method and device

PendingCN121806571AProgramme controlComputer controlText annotationLayout
The embodiment of the invention discloses a machine-readable building engineering standard digital processing method and a machine-readable building engineering standard digital processing device. A specific embodiment of the method comprises the steps of generating text sequence information according to a first preset processing mode; generating layout information and bibliography and reference relation information according to preset layout annotation information, the layout level identification model and preset bibliography annotation information; text information is generated according to preset text labeling information and a paragraph extraction model; according to a pre-trained formula identification model, generating special format text information; determining the text sequence information, the layout information, the bibliography and reference relation information, the text information and the special format text information as construction standard control information; and generating control parameter information according to the construction standard control information, and controlling the construction robot to perform construction processing according to the control parameter information. According to the embodiment, the accuracy of obtaining the building engineering standard information is improved, the construction time consumption is shortened, and the waste of equipment resources is reduced.
Owner:CHINA INST OF BUILDING STANDARD DESIGN & RES

CAD drawing information extraction method for constructing multiple agents based on multi-modal large model

ActiveCN121765486AGeometric CADSemantic analysisSemantic alignmentText annotation
The invention provides a CAD drawing information extraction method for constructing multiple agents based on a multi-modal large model, and belongs to the technical field of multi-modal large models. Geometric figure visual features and text annotation language features are extracted and aligned by using a spiral progressive combination network and a multi-head cross attention cross-modal semantic alignment model, and hierarchical attention processing is performed by using quadtree space division in combination with dense attention and a sparse global token communication mechanism. Initial information extraction and consistency check are carried out by utilizing an analysis agent and a verification agent, when the consistency confidence is lower than a preset threshold value, primitives are re-divided through minimum segmentation optimization, information extraction is carried out, and finally complete drawing information is output. The technical problem of high matching error rate of geometric primitives and text annotations caused by inaccurate cross-modal semantic alignment during CAD drawing information extraction is solved.
Owner:BEIJING NANCAL RUIYUAN DIGITAL TECH CO LTD

Drawing updating identification method and system

PendingCN121640506ACharacter and pattern recognitionGeometric primitiveSoftware engineering
The invention discloses a drawing update identification method and system, and the method comprises the following steps: S1, analyzing drawing files of new and old versions, and extracting multi-dimensional information including geometric primitives, layers and text annotations, S2, comparing the extracted multi-dimensional information of new and old versions, and identifying difference items, S3, based on a preset rule base, analyzing the identified difference items, and carrying out the identification of the new and old versions of the new and old versions of the new and old versions of the new and old versions of the new and old versions. The method comprises the following steps of S1, comparing and analyzing a drawing version, S2, automatically judging the change type and the influence level of the drawing version, S4, automatically generating a visual change report on the basis of a comparison and analysis result, highlighting a change area in a graphical mode by the report, and attaching a structured change list.According to the method, a traditional drawing version management mode is thoroughly innovated through an automatic, multi-dimensional and intelligent technical means; accurate, efficient and intelligent change identification is realized, project risks are effectively reduced, and cooperation efficiency and management level are improved.
Owner:ZHANGJIAKOU POWER SUPPLY COMPANY OF STATE GRID JINBEI ELECTRIC POWER COMPANY

A speech synthesis method based on an implicit continuous consistency model

PendingCN122313940AData setText annotation
This invention relates to the field of speech synthesis technology, specifically to a speech synthesis method based on an implicit continuous consistency model. The method involves constructing a dataset containing audio and its text annotations; building a residual vector quantization variational autoencoder, training it using a joint loss until convergence, and extracting the mean and variance of all audio data mapped to a latent vector distribution; sampling Gaussian noise latent variables based on the mean and variance, sampling Gaussian noise and time steps, and adding noise to obtain speaker features; calculating the continuous consistency loss using the time steps, speaker features, text, audio data, and the consistency model; and optimizing the continuous consistency loss until the consistency model converges. This method utilizes residual vector quantization technology to achieve high-rate audio feature compression and decoupling, and combines the single-step sampling characteristics of the consistency model in the latent space to improve the inference efficiency and training stability of speech synthesis, enabling the model to generate high-fidelity speech in a very small number of iterations.
Owner:HARBIN INST OF TECH AT WEIHAI +1

Data processing method and apparatus, and model optimization method and apparatus

Provided are a data processing method and apparatus, and a model optimization method and apparatus. The data processing method comprises: respectively sampling in a plurality of data sets so as to determine target sample data on the basis of sampling results; inputting the target sample data into a plurality of language models for respectively processing to obtain target prediction data respectively outputted by the plurality of language models; using a preset text annotation model to perform text annotation processing on each piece of target prediction data to obtain text annotation information corresponding to each piece of target prediction data; and constructing a sample data group on the basis of the target sample data, the target prediction data, and the text annotation information, wherein the sample data group is used for executing a model training task and a model verification task.
Owner:INFLY TECH (SHANGHAI) CO LTD

A Controllable Generation Method for Digital Printing Pattern Layout Based on Multi-Factor Attention Excitation

PendingCN122089867AImprove layout control accuracyBiological modelsEditing/combining figures or textTextile printerData set
This invention discloses a controllable generation method for digital printing pattern layout based on multi-factor attention excitation, specifically including the following steps: Step 1, establishing a digital printing pattern dataset containing text annotations, which includes text descriptions, bounding box diagrams, and digital printing patterns; Step 2, training the model on the dataset constructed in Step 1 using the ControlNet framework based on a diffusion model to obtain model training weights; Step 3, using the weights trained in Step 2 for sampling inference, extracting cross-attention weights and self-attention weights during the inference process; Step 4, calculating the cross-attention and self-attention losses inside and outside the bounding box for the attention weights extracted in Step 3; Step 5, updating the noisy image through stepwise loss minimization and gradient descent; Step 6, generating a clear printing pattern through multi-step iterative denoising. This invention solves the problem of inaccurate layout control in existing pattern layout control methods.
Owner:XI'AN POLYTECHNIC UNIVERSITY

Cross-modal medical video segmentation method and system based on text reference

The invention provides a cross-modal medical video segmentation method and system based on text reference, and relates to the technical field of medical video segmentation and computer vision, the cross-modal medical video segmentation method based on text reference comprises the following steps: collecting cross-domain medical videos and text labeling data, and constructing a text-medical video segmentation data set; constructing a cross-modal medical video segmentation model based on text reference, wherein the cross-modal medical video segmentation model comprises a coding module, a cross-modal time sequence context aggregation module and a decoding module; and iteratively training the cross-modal medical video segmentation model through the text-medical video segmentation data set until a training completion condition is reached, and outputting the segmented medical video by the trained cross-modal medical video segmentation model based on the input medical video and text information. According to the segmentation method, representative key frames can be obtained through segmentation, so that a doctor can provide more accurate diagnosis based on a dynamic image.
Owner:WUHAN UNIV OF TECH

Text labeling method, system, device and medium based on multi-model cooperation

PendingCN122309751ATheoretical computer scienceText annotation
This invention relates to the field of artificial intelligence technology, specifically providing a text annotation method, system, device, and medium based on multi-model collaboration. The method includes: acquiring user-submitted text to be annotated, annotation requirements, and performance constraint parameters; parsing the annotation requirements to generate a standardized annotation task set and extracting text features; decomposing the tasks into atomic annotation tasks and constructing task execution paths based on dependencies; using a multi-attribute decision algorithm to match the optimal model for each atomic task based on text features, performance constraints, and a model profile library, and generating a scheduling plan according to the execution path; invoking multi-model collaborative inference to obtain the original annotation results, and outputting structured annotation results after fusion processing. This invention achieves dynamic decoupling between annotation tasks and models, supports multi-model collaborative inference and result verification, and significantly improves the flexibility, accuracy, and efficiency of text annotation.
Owner:浪潮智慧科技有限公司 +2

A hierarchical multi-label attribution method and system fusing atomic rule-driven trustworthy features and knowledge distillation

This invention discloses a hierarchical multi-label attribution method and system that integrates atomic rule-driven credible features and knowledge distillation, belonging to the field of natural language processing technology. First, this invention constructs an atomic rule base for weakly supervised text annotation. Then, it uses a large language model as a teacher model to correct and supplement the weak annotation results, extracting the probability distribution of soft labels and intermediate layer feature representations on each level of labels. Next, it evaluates the credibility of the teacher model's output, selecting a subset of credible soft labels and credible feature dimensions. Then, it constructs a student model with a hierarchical output structure, designs a joint loss function, and distills the student model for training. Finally, it deploys only the student model for inference, outputting hierarchical multi-label attribution results and key evidence fragments. This invention, through the combination of atomic rules and credible knowledge distillation, significantly reduces inference costs while improving the accuracy, stability, and interpretability of hierarchical multi-label attribution.
Owner:THE THIRD RES INST OF MIN OF PUBLIC SECURITY

Power inspection credible detection method and system based on visual language model thinking chain and rule perception reinforcement learning

The invention discloses an electric power inspection credible detection method and system based on a visual language model thinking chain and rule perception reinforcement learning, and relates to the technical field of computer vision, large model application and artificial intelligence safety monitoring. A large visual language model is driven to generate an explicit thinking chain text before outputting a detection box, and the problems that existing detectors such as YOLO / DETR are lack of semantic reasoning ability, black box decision cannot be explained and generalization ability is poor under the small sample condition are solved. According to the method, an end-to-end large visual language model architecture is adopted, a rule perception reinforcement learning mechanism is combined, an electric power safety regulation is converted into a computable logic reward function, and physical constraint and compliance verification are carried out on a reasoning process. According to the method, high-precision detection is guaranteed, semantic-level interpretation of violation behaviors is achieved, the data efficiency and the system credibility in a few-sample scene are remarkably improved, and the method is suitable for power operation safety monitoring.
Owner:HEBEI POWER CONSTR SUPERVISION CO LTD +1

Annotating textual data

PendingFR3170062A1Semantic analysisText annotationSyntax
The invention relates to a method for iteratively annotating textual data implemented in a computer device (CD), said method comprising the following at a current iteration (i): - receiving (S2) an unannotated text (Ti) as input from at least three different text annotation models, the at least three models (MA1, MA2, MA3) having been trained from a first corpus of training data comprising annotated textual data, - generating (S4) three annotated texts respectively (TA1i, TA2i, TA3i), - selecting (S5) one of the three annotated texts, the selected annotated text (TAsi) being the one that contains the fewest semantic and / or syntactic differences, compared with the other two annotated texts, - constructing (S7) a new corpus of annotated textual data (N_CO) by adding to it the selected annotated text (TAsi) and the unannotated text (Ti),said new corpus constituting a second corpus of training data for the text annotation model. Figure 3,
Owner:ORANGE SA

Parameter setting method for training text detection model, training method of text detection model, text detection method, device, equipment and computer program product

The invention discloses a parameter setting method for training a text detection model, a training method of the text detection model, a text detection method, a device, equipment and a computer program product. The parameter setting method comprises the following steps: acquiring an original text image, a text annotation box in the original text image and a dynamic adjustment constraint parameter of a scaling parameter; calculating a current scaling parameter corresponding to the text labeling box according to the dynamic adjustment constraint parameter; shrinking the text labeling box by using the current scaling parameter to obtain a text shrinking box; restoring the text contraction box to obtain a text restoration box; and calculating a final scaling parameter corresponding to the text annotation box according to the text restoration box, the text annotation box and the dynamic adjustment constraint parameter. According to the method and the device, unclipratio is adaptively optimized in the preprocessing stage of model training, so that text regions with different scales and sizes can be dynamically adapted, and the training effect of the model and the text detection precision are further improved.
Owner:RICOH SOFTWARE RES CENT BEIJING

Open vocabulary multi-label image classification method based on hierarchical dual-granularity alignment

The invention relates to an open vocabulary multi-label image classification method based on hierarchical dual-granularity alignment. The open vocabulary multi-label image classification method comprises the following steps: selecting a public image data set with multiple labels; a prompt template is constructed, and text processing is carried out on the image annotations; visual embedding and text embedding are obtained through image and text labeling; respectively obtaining corresponding feature vectors through visual embedding and text embedding; and obtaining a final multi-label classification result of the image by using the feature vector. Compared with traditional multi-label classification, the method has more accurate image classification capability.
Owner:CHONGQING UNIV

A hierarchical multi-label attribution method and system that integrates atomic rule-driven trusted features and knowledge distillation

This invention discloses a hierarchical multi-label attribution method and system that integrates atomic rule-driven credible features and knowledge distillation, belonging to the field of natural language processing technology. First, this invention constructs an atomic rule base for weakly supervised text annotation. Then, it uses a large language model as a teacher model to correct and supplement the weak annotation results, extracting the probability distribution of soft labels and intermediate layer feature representations on each level of labels. Next, it evaluates the credibility of the teacher model's output, selecting a subset of credible soft labels and credible feature dimensions. Then, it constructs a student model with a hierarchical output structure, designs a joint loss function, and distills the student model for training. Finally, it deploys only the student model for inference, outputting hierarchical multi-label attribution results and key evidence fragments. This invention, through the combination of atomic rules and credible knowledge distillation, significantly reduces inference costs while improving the accuracy, stability, and interpretability of hierarchical multi-label attribution.
Owner:THE THIRD RES INST OF MIN OF PUBLIC SECURITY

Intelligent extraction system and method for cargo multimedia information

ActiveCN121526262BImplement automatic conversionreduce distractionsMultimedia data indexingSemantic analysisLogistics managementText annotation
The present application belongs to the technical field of intelligent logistics, and discloses a cargo multimedia information intelligent extraction system and method, which comprises the following steps: collecting cargo pictures, video clips and text annotations uploaded by cargo owners, combining with device positioning signals to generate position coordinate binding, and obtaining a multimedia position set; performing content hierarchical scanning, separating visual elements and dynamic sequences to form a hierarchical structure, and obtaining a hierarchical content group; performing cross semantic bridging, constructing an inter-element correlation path, and obtaining an integrated semantic chain; applying attribute extraction cycles to the integrated semantic chain, extracting cargo specification details and transportation constraint segments to form an attribute set, and obtaining a structured attribute set; generating a driver query sequence based on the structured attribute set and the position coordinate binding, integrating a path order factor for adaptive sorting, and obtaining an optimal driver list; and greatly improving the accuracy of cargo supply and demand matching.
Owner:上海新颐科技软件股份有限公司

Cargo multimedia information intelligent extraction system and method

The invention belongs to the technical field of intelligent logistics, and discloses a cargo multimedia information intelligent extraction system and method, and the method comprises the steps: collecting a cargo picture, a video clip and a text label uploaded by a cargo owner, generating position coordinate binding in combination with an equipment positioning signal, and obtaining a multimedia position set; performing content hierarchical scanning, separating the visual elements and the dynamic sequence to form a hierarchical structure, and obtaining a hierarchical content group; performing cross semantic bridging, and constructing an inter-element association path to obtain an integrated semantic chain; attribute extraction circulation is applied to the integrated semantic chain, cargo specification details and transportation constraint fragments are extracted to establish an attribute set, and a structured attribute set is obtained; based on the binding of the structured attribute set and the position coordinates, generating a driver query sequence, and fusing the driver query sequence into a path-in-the-way factor to carry out adaptive sorting to obtain a preferred driver list; and the freight supply and demand matching accuracy is greatly improved.
Owner:SHANGHAI XINYI TECHNOLOGY SOFTWARE CO LTD

PID paper element intelligent recognition and topological reconstruction method based on visual detection

ActiveCN121281087BCharacter and pattern recognitionAlgorithmText annotation
The present application relates to a kind of PID drawing element intelligent identification and topological reconstruction method based on visual detection, by using multi-modal information collaborative extraction: component detection based on improved YOLOv11 and all text annotations in the drawing are identified using PaddleOCR, pipeline identification algorithm and merging filtering method based on pixel point's probability Hough transform, and T type connecting point is identified and defined as a kind of special topological node, the positioning information acquisition of component and text in PID drawing is realized, and regular expression and Euclidean distance are used to associate component and text, then image element is converted into topological graph model with semantic relationship, realize the end-to-end automation process of "detection-association-reconstruction", finally in the form of image interface display, provide data basis for subsequent drawing analysis, system simulation, equipment management and other applications, end-to-end automatic conversion, greatly improve efficiency.
Owner:CHICHENG TECH

Refining mode in tablet computing device

Embodiments of the present invention provide a computerized system and method in an electronic paper tablet device to provide a user with two modes for interacting with the electronic paper tablet device. In the authoring mode, the user can input text, make annotations, draw and take other authoring actions. In the refinement mode, the user can edit texts, annotations and drawings and perform other refinement on actions taken by himself / herself in the creation mode. Both hardware and software of the electronic paper tablet support both modes, including special hardware keys on a keyboard associated with the electronic paper tablet that enable a user to switch between the two modes. Thus, embodiments of the present invention enable a user to closely process text, edit and comment text using a stylus device (or even an input mechanism such as a finger), while maintaining integrity with underlying text and its display.
Owner:REMARKABLE AS

Data augmentation method, device and system based on combination of three-dimensional scene and language data

The application provides a data enhancement method, device and system based on three-dimensional scene and language data combination, the method comprising: obtaining 3D scene data and corresponding text annotation data; preprocessing the scene data and the text annotation data respectively to obtain preprocessed 3D-language combined data; sequentially performing multi-modal data enhancement and semantic quality filtering processing on the preprocessed 3D-language combined data to obtain a target 3D-language combined data set. The application integrates various data sources such as 3D point cloud data, RGB-D images, question and answer pairs and dense descriptions, uses data preprocessing, multi-modal data enhancement and semantic quality filtering to realize automatic construction of a high-quality large-scale data set, can improve the data quality of 3D scene understanding and visual question and answer tasks while enhancing the diversity and generalization ability of the data, and provides strong support for 3D visual understanding, robot task planning and other applications.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Method and device for processing dialogue text, electronic equipment and storage medium

ActiveCN115563255BText recognitionText annotation
This invention provides a method, apparatus, electronic device, and storage medium for processing dialogue text. In response to a text recognition operation on the dialogue text, the method acquires the dialogue text content to be recognized, then acquires the dialogue text model corresponding to the dialogue text content, and obtains the target dialogue role features corresponding to the dialogue text content from the dialogue text model. Using the target dialogue role features as a text index, the method determines the storage location of the dialogue text content in the dialogue text model, then acquires at least one predicted class label corresponding to the storage location, decodes the predicted class label to obtain the text annotation object corresponding to the predicted class label, and extracts the content corresponding to the text annotation object. This improves the recognition accuracy of text information and further enhances the accuracy of extracting important information from the dialogue text.
Owner:BEIJING SINOVOICE TECH CO LTD

Search engine knowledge index generation method and related equipment

The invention relates to the field of AI. The invention particularly relates to a search engine knowledge index generation method and related equipment. The method comprises the following steps: segmenting an original text to obtain a plurality of text blocks; generating text annotation information of each text block according to each text block in the plurality of text blocks; and / or, context information of each text block is generated according to each text block, and the context information of each text block comprises basic context information of each text block and an identifier of an associated text block of each text block; and the index generation device generates an index of each text block according to the comment information of each text block and / or the context information of each text block. By adopting the scheme provided by the invention, the accuracy of a retrieval result is improved.
Owner:HUAWEI TECH CO LTD

Multi-modal large model driven geometric image reading and analyzing method and system

The invention discloses a multimodal large model driven geometric image reading analysis method and system, and relates to the technical field of data processing, the method comprises the following steps: after receiving image data uploaded by a user, extracting a geometric element candidate set, and setting an initial confidence coefficient for each candidate element; starting a geometric annotation recognition channel in parallel, extracting a geometric element candidate set by using a visual encoder, generating visual embedding, activating an OCR channel to perform text recognition, and generating text semantic embedding; and generating structured geometric representation through confidence weighted cross-modal alignment fusion, inputting the structured geometric representation into the hybrid reasoning recognition model, and outputting an analysis result. The technical problems that an existing geometric image analysis method cannot effectively process image and text information at the same time, and the analysis precision and efficiency of a complex geometric structure are low are solved, and the technical effects that the geometric information and text labels in the image are accurately extracted through fusion of the multi-modal large model, and the analysis precision and efficiency are improved are achieved.
Owner:JIANGSU HAOHAN INFORMATION TECH +1

Fraud confrontation dialogue generation and detection method and system based on large model and multi-agent

PendingCN122157654ANatural language data processingSpeech recognitionText annotationModal voice
The application provides a fraud confrontation dialogue generation and detection method and system based on a large model and multiple agents, wherein the method comprises the following steps: scene construction; multi-modal voice data generation; text annotation based on slow thinking and multi-round deep reasoning analysis. The application proposes a fraud confrontation dialogue generation and deep thinking analysis scheme that fuses a large language model and multiple agents, constructs a new generation of anti-fraud intelligent system with dynamic attack and defense game capabilities, constructs a specialized agent cluster, fills the key research gap in the field of multi-modal fraud detection, solves key problems such as data privacy and scene diversity, and provides technical support for promoting the development of an intelligent anti-fraud system. The application deeply fuses voice data generation technology, greatly improves the complexity and practicality of confrontation training, and approximates a real fraud scene.
Owner:NANJING YAXIN SOFTWARE CO LTD