Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

226 results about "Lexical analysis" patented technology

In computer science, lexical analysis, lexing or tokenization is the process of converting a sequence of characters (such as in a computer program or web page) into a sequence of tokens (strings with an assigned and thus identified meaning). A program that performs lexical analysis may be termed a lexer, tokenizer, or scanner, though scanner is also a term for the first stage of a lexer. A lexer is generally combined with a parser, which together analyze the syntax of programming languages, web pages, and so forth.

Zero-illusion large language model output method and system

The invention relates to the technical field of large language models, and discloses a zero-illusion large language model output method and system, and the method comprises the steps: sequentially carrying out the deterministic lexical analysis and standardized mapping of a received original query, and generating a standard term set; locking to-be-reasoned items through hierarchical matching to form a deterministic retrieval result set; executing five rounds of progressive verification processing on the result set to obtain an integrated result set with an LLM verification mark; and outputting a layered zero illusion result through multiple verification and auditing of structure compliance, semantic consistency and knowledge base double recheck. According to the method, probabilistic reasoning is replaced by full-link deterministic operation, so that each code can be traced back to a credible knowledge source, illusion is thoroughly eliminated, absolute reliability and complete traceability of an output result in a professional field are realized by a verifiable deterministic path, and a foundation is laid for zero-illusion application.
Owner:YUANYUANZHIJI ARTIFICIAL INTELLIGENCE TECHNOLOGY (CHONGQING) CO LTD

ANTLR4-based multi-manufacturer network equipment configuration analysis method and system

The invention discloses an ANTLR4-based multi-manufacturer network equipment configuration analysis method and system, and the method comprises the steps: carrying out the preprocessing of a configuration file; based on a machine learning algorithm of multi-dimensional feature analysis, automatic identification of configuration file manufacturer types is realized; based on the identification result of the manufacturer type, loading an ANTLR4 grammar parser corresponding to the manufacturer type; a lexical analyzer in ANTLR4 realizes lexical analysis on the configuration file through lexical rule definition, a Token classification system and an optimized analysis algorithm; a syntactic analyzer in ANTLR4 constructs an abstract syntax tree according to a lexical analysis result and a syntactic rule defined based on a context-independent grammar; traversing the abstract syntax tree, and extracting key configuration information; and converting the specific configuration format of each manufacturer into a unified standard model. According to the method and the system, network equipment configuration files of different manufacturers can be automatically identified and analyzed, and key configuration information is extracted and converted into a unified model.
Owner:CHINA UNITECHS

Drug design method based on autoregressive model

A drug design method based on an autoregressive model is provided, which relates to the field of drug design technologies. The method includes: applying a sub-word tokenization algorithm to biological text processing, training protein and ligand information in data sets to obtain a protein tokenizer and a ligand tokenizer, and constructing a tokenizer of the autoregressive model; processing and transforming original data in the data sets into a text form, and encoding by the tokenizer to construct a training data set for the autoregressive model; training the autoregressive model by the training data set, so that the autoregressive model can understand SMILES representations of ligands and learn an interaction mode between proteins and ligands; generating predicted ligands by using the trained autoregressive model, and post-processing through a chemical information tool to acquire candidate ligands with specific chemical structures; and evaluating and optimizing the candidate ligands to determine target candidate molecules.
Owner:THE FIRST AFFILIATED HOSPITAL OF MEDICAL COLLEGE OF XIAN JIAOTONG UNIV

Code error repairing method and device based on compiler

PendingCN120743278ACode compilationSemantic propertySyntax error
The embodiment of the invention provides a code error repairing method based on a compiler, which comprises the following steps: receiving a target code written by using a target language, and executing a first repairing operation on the target code in the process of compiling the target code by the compiler to obtain a first repairing sequence, the first repairing operation comprises lexical error repairing corresponding to the lexical analysis process and grammatical error repairing corresponding to the grammatical analysis process. Performing a plurality of rounds of semantic repair operation based on an abstract syntax tree, wherein the abstract syntax tree used for the first round of semantic repair operation is generated after the first repair sequence is subjected to syntactic analysis by the compiler; any round of semantic repair operation comprises the steps of cloning a current abstract syntax tree to obtain a cloned syntax tree, and performing semantic attribute labeling based on the cloned syntax tree; and for any error node with the error attribute label in the clone syntax tree, performing semantic repair on a node corresponding to the abstract syntax tree.
Owner:ALIPAY (HANGZHOU) INFORMATION TECH CO LTD

Latency-based resource allocation in model-as-a-service platform

A model-as-a-service (MaaS) platform performs cross-model resources allocation from a shared pool of GPU resources based on model-agnostic metrics generated by a metric standardizer. The metric standardizer receives, from model providers, model-specific benchmark metrics that define relationships between resource utilization and token processing according to the different model-specific tokenization schemes; receives, from one or more MaaS components, token-based job metrics pertaining to LLM processing tasks; and determines, based on the model-specific benchmark metrics and token-based job metrics, the model-agnostic metrics for multiple model pools executing instances of different large language models (LLMs) that generate and process text according to different model-specific tokenization schemes. The MaaS platform further includes one or more resource allocation components that dynamically reallocates resources of the shared pool based on the model-agnostic metric.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Transverse mixed attention mechanism model training method, medium, device and program product

The invention provides a model training method for a transverse mixed attention mechanism, a medium, equipment and a program product, and the method comprises the steps: obtaining a data set containing a plurality of sample sequences, each sample sequence in the data set being formed by arranging a plurality of Token sequences obtained through word segmentation; constructing a to-be-trained model based on the pre-trained full attention model, and adding newly added parameters for linear attention calculation; in the same transverse mixed attention layer, executing total attention calculation on a Token set in a preset total attention calculation range, executing linear attention calculation on all Tokens, and fusing results of the total attention calculation and the linear attention calculation to obtain transverse mixed attention output used for forward reasoning and loss calculation; and based on the output and prediction result, only updating the newly added parameters to optimize the to-be-trained model until the to-be-trained model converges. According to the method, the calculation complexity and video memory occupation of long text sequence processing are reduced, and the reasoning speed and the resource utilization rate are improved.
Owner:BEIJING JIBU QIANLI TECHNOLOGY CO LTD

Hybrid recommendation system using Team Frequency-Inverse Document Frequency (TF-IDF) for film recommendations

A hybrid recommendation system for generating personalized film recommendations, the system comprising the following: a processing unit for receiving metadata, configured to receive descriptive data related to films, the descriptive data including plot summaries, genre identifiers, cast lists and director names; a text preprocessing module that is communicatively coupled with the metadata ingestion processing unit and is configured to perform tokenization, stopword removal, stemming, and lemmatization on the descriptive data to obtain a processed corpus; a TF-IDF vectorization unit configured to encode the processed corpus into high-dimensional semantic feature vectors by calculating TF-IDF inverse document frequency values ​​over the entire film dataset, with the semantic feature vectors representing the contextual meaning of terms for individual films; a collaborative filter engine comprising a user-element interaction matrix, wherein the engine is configured to compute latent preference signals using one or more techniques selected from the group consisting of cosine similarity, k-nearest neighbor similarity, and matrix factorization; a score fusion controller coupled to both the TF-IDF vectorization unit and the collaborative filter engine, wherein the module is configured to normalize the semantic feature vectors and the collaborative preference scores and dynamically combine them according to an adaptive weighting coefficient, the coefficient being determined as a function of data density, interaction sparsity, and user history length; and a recommendation output module configured to evaluate candidate films for each user based on the combined score and generate a top-N recommendation list.
Owner:NITTE MEENAKSHI INSTITUTE OF TECHNOLOGY (DEEMED TO BE UNIVERSITY) BENGALURU +3

Abnormal access detection method and system based on data analysis

The invention aims to provide an abnormal access detection method and system based on data analysis, and relates to the technical field of network data communication, user behavior classification is carried out based on a support vector machine, only normal behavior data is needed to establish a detection model, abnormal data is not needed, the model training complexity is highly simplified, and the user experience is improved. Compared with a traditional method, the OCSVM has distinct advantages in the aspects of single-class data processing and anomaly detection; according to the method, log data are analyzed by using an LEX lexical analyzer and a YACC syntactic analyzer, high-precision description of user behaviors is realized, feature selection is performed on original logs, sensitive information such as user names, operation time and client IPs is extracted, data redundancy is reduced, and high efficiency of model training and recognition is ensured; according to the anomaly detection method provided by the invention, the user behavior anomaly information in the database can be efficiently analyzed, and an alarm can be given in time and necessary countermeasures can be taken, so that potential safety risks or data leakage can be prevented.
Owner:GUANGXI POWER GRID CORP

Method and system for analyzing and editing view in real time based on streaming data

The invention discloses a real-time analysis and view editing method and system based on streaming data, and the method comprises the steps: receiving streaming data which comes from a large language model and is organized in an editing protocol format, carrying out the real-time lexical analysis of the streaming data, and decomposing the streaming data into a semantic unit sequence through a semantic pattern recognition technology; performing grammar analysis on the semantic unit sequence, and converting the semantic unit into a structured editing instruction for modifying an existing rendering tree node according to an instruction pattern recognition mechanism and a target object analysis mechanism; converting the structured editing instruction into an atomic instruction sequence executable by a rendering tree editing interface; and modifying the rendering tree in real time according to the atomic instruction sequence and triggering update rendering of the view. According to the method, the content generated by the AI is analyzed in real time by adopting a streaming output mechanism, so that a user can observe the generation process of the view in real time, local redrawing of the view is realized by using an editing protocol, and the view editing efficiency is greatly improved.
Owner:HANGZHOU DIMENG TECHNOLOGY CO LTD

Nuclear power equipment operation defect grading method and device, electronic equipment and storage medium

The invention relates to the technical field of nuclear power plants, in particular to a nuclear power equipment operation defect grading method and device, electronic equipment and a storage medium. The invention provides a nuclear power equipment operation defect grading method, and aims to realize grading of nuclear power equipment operation defects through a systematic and multi-level fusion decision-making mechanism. The problems that rule grading lacks flexibility, graph grading is influenced by keyword sensitivity, large model grading field knowledge is insufficient, and precision and adaptability are limited due to the fact that a traditional fusion strategy is single in the prior art are effectively solved. The core of the method is that lexical meta processing and a value matrix mechanism are introduced, dynamic and fine-grained fusion of multi-base classifier output is achieved, and therefore the overall performance of a grading system is improved. Through lexical element processing, value matrix construction and dynamic updating and multi-classifier fusion decision making, the nuclear power equipment defect grading method with high precision, high adaptability and high reliability is realized, and many limitations in related technologies are effectively solved.
Owner:CHINA NUCLEAR POWER ENGINEERING COMPANY LTD +1

Code similarity detection method and system based on large language model

The invention discloses a code similarity detection method and system based on a large language model, and the method comprises the steps: obtaining a to-be-detected source code pair, and marking the to-be-detected source code pair as a source code A and a source code B; performing code cleaning and format standardization on the source code A and the source code B, and mapping a variable name and a function name which are customized by a user into a uniform placeholder; analyzing the source codes A and B based on the abstract syntax tree, respectively replacing variable names and function names in the source codes A and B with unified serialized placeholders, and maintaining a mapping table; meanwhile, expanding a lexical dictionary of the pre-training large language model, and inserting a special identifier; splicing the replaced code snippets with special identifiers, and constructing a structure sensing input sequence; inputting the structure perception input sequence into a pre-trained large language model backbone network for feature coding to obtain a high-dimensional semantic feature vector containing global context information; connecting a multi-task prediction head behind the large language model backbone network, inputting the feature vector into the multi-task prediction head, outputting probability distribution of code clone types through a classification task head, and respectively outputting a row level similarity score and a lexical element level similarity score through a regression task head; and according to the classification probability and the regression score, performing comprehensive judgment by combining a preset threshold, and generating a detection report. According to the method, similar codes after variable renaming, statement rearrangement or control flow transformation can be accurately recognized, and the accuracy and robustness of code similarity detection are improved.
Owner:NANJING UNIV OF SCI & TECH

Query statement display method and device, computer equipment, readable storage medium and program product

The invention relates to a query statement display method and device, computer equipment, a readable storage medium and a program product. The method comprises the steps that in the process that a user inputs an initial query statement, a current character input by the user is scanned in real time based on a lexical analyzer in a regular matching mode; switching a dynamic lexical analysis mode of a lexical analyzer based on the character type of the current character, and analyzing the current character based on the lexical analyzer after mode switching; based on the analysis result of the current character, generating a structured lexical unit Token stream of the initial query statement, and based on a current Token corresponding to the current character in the structured Token stream, obtaining at least one recommendation value of the current character; and generating a target query statement based on a selection result of the recommended value by the user, and displaying the target query statement in a display interface of the user. According to the method provided by the invention, the actual input requirement of the user can be met.
Owner:CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1

General data exchange system based on configurable label structure

The invention discloses a general data exchange system based on a configurable label structure, which relates to the technical field of data processing and information integration, and comprises a parameter module used for establishing an identity authentication mechanism of two exchange parties; the business document module is used for establishing a data structure definition based on a preset exchange demand; the receipt logic module is used for carrying out logic analysis and code decoupling on the data structure definition, generating a marked language script code and an execution instruction, constructing an execution matrix and realizing mapping between a database physical field and a memory logic variable; and the standard adaptation module is used for comparing and checking the content of the generated data and feeding back a checking result to the parameter module to correct the exchange rule and parameter configuration. A more intelligent logic separation strategy is implemented through structural insight, data reading, writing and transmission are cooperatively controlled, complete decoupling of service logic and program codes is achieved through a nested tag structure, and flexibility, universality and maintenance efficiency of data exchange are optimized.
Owner:BEIJING LIGONGDAXUE PRESS CO LTD

Inline Nested Data Loss Protection (DLP)

The disclosure presents systems and methods for hierarchical classification of input data across a plurality of categories. A machine learning model processes various data formats, starting with dimensional reduction using tokenization techniques, such as Bert-tiny tokenization, to create model-readable representations. The system predicts super-categories, sub-categories, and granular categories through selective activation of sub-layers tied to identified super-categories, optimizing computational efficiency. Label smoothing during training mitigates overconfidence in predictions, while softmax normalization refines inference outputs. Synthetic data generation using Large Language Models (LLMs) supplements training datasets, and an automated data labeling pipeline efficiently generates hierarchical labels. Modifications to the model, such as stop word removal and file size limitations, further reduce latency. Inference analyzes logits to predict hierarchical paths, providing detailed classifications with clear outputs. The method is adaptable for multimodal formats, ensuring scalable and accurate predictions across diverse data types while minimizing computational costs and improving reliability.
Owner:ZSCALER INC

Language specification and conversion method and system fused with multi-database dialects

The invention provides a multi-database dialect fused language specification, a multi-database dialect fused language conversion method and a multi-database dialect fused language conversion system, and relates to the technical field of databases. A user inputs a unified SQL statement and selects the type of a target database, lexical analysis and grammatical analysis are performed on the unified SQL statement, an abstract syntax tree is constructed, semantic analysis is performed, dialect characteristics in the unified SQL statement are identified, the support condition of the dialect characteristics in the target database is verified, and a dialect conversion strategy is provided. And selecting a corresponding conversion rule in the unified SQL specification library, converting the grammar of the unified SQL statement into the dialect grammar of the target database, and generating the SQL statement. According to the method, one set of SQL and multi-database support is realized through unified SQL specification formulation, intelligent dialect adaptation and configuration driving, a complete technical solution is provided for cross-database application development, and technical challenges caused by database dialect differences are solved.
Owner:DIGITAL CHINA FINANCIAL SOFTWARE LTD

Utilization-based resource allocation in model-as-a-service platform

A model-as-a-service (MaaS) platform performs cross-model resources allocation from a shared pool of GPU resources based on model-agnostic metrics generated by a metric standardizer. The metric standardizer receives, from model providers, model-specific benchmark metrics that define relationships between resource utilization and token processing according to the different model-specific tokenization schemes; receives, from one or more MaaS components, token-based job metrics pertaining to LLM processing tasks; and determines, based on the model-specific benchmark metrics and token-based job metrics, the model-agnostic metrics for multiple model pools executing instances of different large language models (LLMs) that generate and process text according to different model-specific tokenization schemes. The MaaS platform further includes one or more resource allocation components that dynamically reallocates resources of the shared pool based on the model-agnostic metric.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Source code obfuscation method and system

The invention discloses a source code obfuscation method and system, and belongs to the technical field of software development. According to the source code obfuscation method, in the preprocessing obfuscation stage, all first compiling units in a source code to be obfuscated are subjected to lexical analysis one by one to obtain corresponding lexical units, and a preprocessing instruction tree is generated according to all the lexical units; confusion is carried out in a corresponding mode based on the category of each preprocessing node in the preprocessing instruction tree, so that the purposes of deleting the to-be-confused source code annotation and expanding an internal header file of the to-be-confused source code annotation are achieved; in the symbol name obfuscation stage, all second compiling units in a first source code are analyzed one by one to obtain an abstract syntax tree, a symbol table is constructed according to the abstract syntax tree, the first source code is compiled based on the symbol table to obtain an obfuscated code, and the purpose of renaming information such as variables, functions and types in the source code is achieved; according to the method, the difficulty of quickly analyzing the source code is improved, and the protection strength of the source code is improved.
Owner:ISOFT INFRASTRUCTURE SOFTWARE

General SQL grammar conversion method and device, equipment and medium

The invention relates to the technical field of SQL (Structured Query Language) grammar conversion, and provides a general SQL grammar conversion method, device and equipment and a medium, the method comprises the following steps: performing lexical analysis on a source end SQL statement to obtain a lexical unit set; based on the type features and semantic attributes of the lexical unit set, determining parser combiners, and constructing a specific parser by infinitely combining the parser combiners; the specific resolver is used for conducting grammar analysis on the lexical unit set, a general abstract syntax tree is constructed through a mapping assembly, and nodes of the general abstract syntax tree correspond to the specific resolver one to one; and traversing the general abstract syntax tree, and generating a target end SQL statement according to an SQL syntax rule of a target end database. Through the technical scheme, any complex source end SQL can be accurately converted into the target end SQL.
Owner:CSC FINANCIAL CO LTD

Compilation method for compiling C language source code into RISC-V assembly code

The invention discloses a compiling method for compiling a C language source code into an RISC-V assembly code. The method comprises the following steps: acquiring a C language source code; performing lexical analysis on the C language source code to generate a mark flow; performing syntactic analysis on the mark flow, and constructing an abstract syntax tree; performing semantic analysis on the abstract syntax tree to generate a target abstract syntax tree; constructing a runtime environment of the RISC-V assembly code; generating an intermediate code and a control flow diagram corresponding to the intermediate code according to a rule in the runtime environment and the target abstract syntax tree; generating a target code by using the intermediate code and the control flow diagram corresponding to the intermediate code; according to the technical scheme, the intermediate representation more adaptive to RISC-V custom instruction mapping can be generated, so that the execution efficiency of assembly codes is improved; and meanwhile, by designing a lightweight runtime environment, the performance overhead is further reduced.
Owner:CHINA SOUTHERN POWER GRID COMPANY

System for hardware-based compliance with legal regulations in blockchain smart contracts

A system for autonomous compliance with legal regulations in smart contracts; the system includes: a regulatory data collection unit comprising a network interface controller physically connected to an external communications port, a hardware-based public key infrastructure circuit configured to validate digital certificates of regulatory servers, and a direct memory access controller configured to transfer authenticated regulatory update packets from the network interface controller to a volatile buffer memory without processor intervention; a Compliance Code Conversion Unit with a hardware lexicon scanner implemented as a finite state machine and embedded in reconfigurable FPGA logic blocks, a microcontroller executing a hardware-based lexical analysis pipeline stored in firmware registers, and a translation cache memory configured to temporarily store tokenized compliance rules before writing them to a non-volatile instruction memory; a policy evaluation and enforcement unit comprising a transaction verification processor connected to a secure enclave memory, a hardware comparator circuit configured to compare transaction parameters with compliance thresholds stored in the secure enclave memory, and a logic gate circuit configured to generate execution gate signals that selectively allow, block, or modify smart contract execution signals transmitted to a blockchain execution processor; and an audit logging unit with a cryptographic hashing circuit configured to generate block hashes of enforcement records, a timestamp oscillator configured to generate temporal signatures of compliance enforcement actions, and a Merkle tree generation circuit configured to create tamper-proof hierarchical hash structures, with the enforcement records being passed to a decentralized storage interface for anchoring in the chain or distributed ledger.
Owner:KEMPAIAH MADHURA GAYATHRI BENGALURU +3

Large language model operator compiling method and system oriented to GPU (Graphics Processing Unit) platform

The invention provides a large language model operator compiling method oriented to a GPU platform, and belongs to the field of GPU operator compilation. A large language model operator is compiled through a Python interface to obtain a Python code; performing lexical analysis and grammatical analysis on the Python code to generate an abstract syntax tree; mapping the abstract syntax tree into an MMIR extension intermediate representation; tensor layout in the MLIR extension intermediate representation is uniformly converted into a binary linear mapping matrix, and optimization problems of global memory layout, shared memory layout and register block layout are analyzed based on the binary linear mapping matrix; automatically deducing a non-key tensor layout based on a program control flow diagram, generating a global optimal layout scheme, and obtaining an optimized MMIR extension intermediate representation; enabling the optimized MLIR extension intermediate representation to pass through a GGPU (Graphics Graphics Processing Unit) target code generator of the MLIR, and generating a PTX code adaptive to the NVIDIA GPU; the PTX code is loaded to a GPU, and an operator calculation result is output; the invention further provides a compiling system. The problem that the current large language model operator compiling efficiency is low is solved.
Owner:江淮前沿技术协同创新中心

Method and system for dynamic weighted metrics-based evaluation and tokenization of large language models

The embodiments of the present disclosure herein address unresolved problems of evaluation of LLM response quality and overall LLM models. Existing approaches for LLM evaluation and LLM response evaluation can be broadly categorized into automatic evaluation metrics, human evaluation, and adversarial testing. Embodiments herein provides a method and system for dynamically weighted selection of performance metrics for generation of LLM response score. Further, the system is configured method and system for generation of LLM maturity gap analysis and associated recommendation for improvement of LLM response score. Finally, the system generates a compliance certificate for every model (version) with a (threshold) level score and generates an NFT using a smart contract based blockchain, using metadata associated with the model and the evaluation metrics and results.
Owner:TATA CONSULTANCY SERVICES LTD

Multimodal automatic driving training method based on DeepSeek training framework

The invention relates to the technical field of automatic driving, in particular to a multi-mode automatic driving training method based on a DeepSeek training framework. Comprising the following steps: reading multi-view camera images and text instructions of a DriveLM-nuScenes data set, and splicing the images according to a look-around layout to form panoramic representation; performing zooming, normalization and standardization processing on the panoramic image to obtain an image tensor; performing marking processing on the text instruction, inserting an image placeholder and a dialogue role mark, and structuring text input representation; dimensionality alignment, position code addition and cross-modal attention fusion of vision and text marking sequences are realized through a multi-modal alignment module, and multi-modal embedding representation is generated; and inputting the embedded representation into a DeepSeek language model to generate a decision text through autoregression, and taking the cross entropy loss with a mask as an optimization target. According to the method, the problems of insufficient multi-view fusion, weak modal alignment and the like in the prior art are solved, the cognitive reliability and the decision interpretability in a complex scene are improved, and vehicle-mounted edge deployment is adapted.
Owner:HEFEI UNIV OF TECH

Method of training language model for cybersecurity and system performing the same

Provided is a system for training a language model for cybersecurity, which includes: a document collection unit that collects a cybersecurity document used for training a language model for cybersecurity; an extraction unit that identifies non-linguistic elements in the cybersecurity document based on a non-linguistic element database; a tokenization unit that tokenizes the cybersecurity document to generate a plurality of tokens; and a language model application unit that controls the language model to simultaneously perform a first task of classifying types of the non-linguistic elements including at least one of a Bitcoin address, a hash value, an IP address, and a vulnerability identifier included in the cybersecurity document and a second task of recovering only linguistic elements of the cybersecurity document.
Owner:S2W INC +1

Automatic configuration method and system for DCS data acquisition interface of power plant

The invention provides a power plant DCS data acquisition interface automatic configuration method and system. The method comprises the following steps: acquiring a production process diagram file from a DCS upper computer, and analyzing and extracting measuring point information based on a DCS type; generating a structured lexicographical order through lexical analysis, and screening target measurement points according to a preset rule to generate a measurement point list; adopting a Myers difference algorithm to compare the states of the measuring point lists, and dynamically updating the interface program acquisition range; and transmitting the measuring point list, the process diagram and the real-time data to a management information layer MIS across a safety area, and synchronously checking configuration consistency based on a digital twinborn model. According to the method, automatic extraction of the measuring points is realized through a regular expression and a DOM analysis technology, dynamic key encryption and an incremental updating mechanism are combined, the problems of low manual configuration efficiency, updating lag and cross-regional data islands are solved, the accuracy and synchronous real-time performance of measuring point configuration are remarkably improved, and the method is suitable for large-scale popularization and application. And safe and reliable digital support is provided for power plant production monitoring and information management.
Owner:HUADIAN ZOUXIAN POWER GENERATION CO LTD +2

Log parsing method and device

Embodiments of the present disclosure provide a log parsing method and device. The log parsing method comprises: determining a log to be parsed, and carrying out tokenization processing on said log to obtain a log tokenization sequence; matching the log tokenization sequence with nodes in a prefix tree, and determining a candidate syntax template corresponding to the log tokenization sequence, wherein the prefix tree is constructed by the nodes and paths, the nodes are tokens in a log syntax template, each path represents an association relationship of tokens in the log syntax template, and the log syntax template is determined on the basis of an initial log; matching the log tokenization sequence with the candidate syntax template on the basis of a preset matching policy to obtain a matching result; and when the matching result does not satisfy a preset matching condition, using a large language model to parse said log, to obtain a parsing result. The method can improve the log parsing accuracy and reduce log parsing costs.
Owner:CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD

Apparatus and methods for a large language model with semantic audio for targeted advertising video stream

A method for selecting and inserting contextually relevant advertisements into a video stream, executed by a processing system in a network server computing device, encompasses receiving a primary video stream with potential advertisement insertion points indicated by SCTE35 / SCTE104 markers, extracting an audio segment from this stream before an advertisement break, obtaining audio from potential advertisements, transcribing both primary and secondary audio segments into textual data, performing semantic analysis and tokenization on this data, creating vector embeddings, and normalizing these embeddings for a feed-forward neural network. The method further involves determining semantic similarity scores between the primary and secondary content through a transformer-based AI model, generating a similarity matrix from these scores, identifying the most contextually aligned advertisement based on these scores, and inserting this advertisement at an indicated break point in the primary video stream.
Owner:CHARTER COMM OPERATING LLC

Method and system for detecting unauthorized URL (Uniform Resource Locator) of website based on automatic operation of common user

The invention provides a method and system for detecting an unauthorized URL of a website based on automatic operation of a common user, belongs to the technical field of network security, and can at least partially solve the problems of low detection efficiency, poor accuracy and excessive dependence on manual operation in the prior art. Automatically logging in a website by using the registration information, processing a verification link, and obtaining session information; obtaining a JS file through a web crawler in a normal user login state; analyzing the JS file, extracting static, dynamic and relative URLs in combination with lexical analysis, grammatical analysis and sandbox simulation execution, and unifying the static, dynamic and relative URLs as absolute URLs; de-duplicating and normalizing the URL; accessing the URL through a common user session by using a network test tool, obtaining response data, comparing the response data with a high-authority user reference response, and judging an unauthorized URL; the method has the beneficial effects that automatic and comprehensive detection of the unauthorized URL is realized, and the detection efficiency and accuracy are remarkably improved.
Owner:HUANENG POWER INT INC +1

Fault-tolerant processing method and system for long text in large model service, and storage medium

The application relates to the technical field of large language model application, in particular to a fault-tolerant processing method and system for super-long text in a large model service and a storage medium. The method comprises the following steps: obtaining a to-be-processed super-long text and a maximum Token threshold of a large model context window, completing Tokenization and judging whether the Token is out of limit; starting a multi-level alternative solution containing fast compression, dynamic compression and chat history compression; starting an error recovery mechanism containing hierarchical exception processing, intelligent degradation and parameter verification; outputting a target text with a compliant Token number; the system comprises a text acquisition module, a multi-level alternative solution execution module, an error recovery module and an output module; the storage medium stores a computer program, and the computer program realizes the above method when executed, can be adapted to multiple models, meets the industrial super-long text processing demand, and solves the problems of the absence of a fault-tolerant mechanism, Token out-of-limit errors caused by insufficient boundary processing and unstable reasoning service in the prior art.
Owner:POWERCHINA BEIJING ENG CORP

A method for identifying data flows outside components based on static analysis

The present invention discloses a method for identifying component external data flows based on static analysis, comprising: reading a component source code file, performing lexical analysis and grammatical analysis on the component source code, obtaining token information contained in the component source code and generating an abstract syntax tree, and obtaining function properties in the component source code based on the information, including: whether it is a function defined in the component source code, whether it is a function imported from outside the component, and whether it is a function exported to outside the component; traversing the abstract syntax tree corresponding to the locally defined function, and maintaining the classification and pointing information of the nodes in the process; for statements containing data update operations, judging the data classification of the data access target, and if it is external component data, adding the source code position corresponding to the statement to the source code position set containing the external component data flow; and finally returning the source code position set containing the external component data flow. The present invention can identify the read and write operations on the external component data flow in the component source code to assist in code partitioning.
Owner:ZHEJIANG UNIV