Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

97 results about "Frame based" patented technology

Frame-based terminology is a cognitive approach to terminology developed by Pamela Faber and colleagues at the University of Granada.One of its basic premises is that the conceptualization of any specialized domain is goal-oriented, and depends to a certain degree on the task to be accomplished. Since a major problem in modeling any domain is the fact that languages can reflect different ...

Semantic guided efficient perspective view to BEV projection and sampling

A device for processing frame data may be configured to identify one or more semantic characteristics for the frame data, wherein the frame data comprises a plurality of frames, each frame of the plurality of frames being acquired for a same scene and by a different frame source of a plurality of frame sources; extract features from each respective frame of the plurality of frames; determine a non-uniform sampling pattern of the plurality of frames based on the one or more semantic characteristics; and project, using the non-uniform sampling pattern, a portion of the extracted features into a bird's-eye-view (BEV) space having a grid structure to generate a fused set of BEV features.
Owner:QUALCOMM INC

Multi-mode large model video content understanding reasoning acceleration method and system

The invention discloses a multi-mode large model video content understanding reasoning acceleration method and system, and mainly relates to the technical field of artificial intelligence reasoning acceleration. Comprising the following steps: inputting video data and preprocessing the video data to generate a video frame sequence; performing adaptive video Token compression on the generated video frame sequence, and outputting a compressed visual Token set; performing visual feature coding and Key-Value generation on the compressed visual Token set to obtain visual KV data; performing video KV cache partition management on the visual KV data; executing cross-modal reasoning based on the vLLM framework to generate a video content understanding result; and outputting a video content understanding result, and carrying out post-processing and structured mapping. The method has the beneficial effects that the obvious reasoning acceleration and throughput improvement can be realized on the premise of keeping the precision of the original large model.
Owner:海看网络科技(山东)股份有限公司

Smart frame selection via activity-based ranking and optimization

Various examples, systems, and methods are disclosed relating to frame selection via activity-based ranking and optimization. A first computing system can receive a plurality of frames and metadata from a capture device capturing a video stream. The first computing system can generate, using a ranking model, a plurality of rankings for the plurality of frames based on a plurality of video parameters of the plurality of frames and the metadata, wherein the plurality of rankings correspond to a summarization of the video stream. The first computing system can determine at least one of the plurality of frames to provide to at least one buffer based on the plurality of rankings, wherein the at least one buffer stores a subset of frames of the plurality of frames. The first computing system can provide, from the at least one buffer, the subset of frames as input to a machine-learning model.
Owner:NVIDIA CORP

Real-time translation method and interaction system based on streaming voice segmentation and semantic verification

PendingCN122050365ANatural language translationSpeech recognitionSpeech segmentationFrame based
The invention provides a real-time translation method and interaction system based on streaming voice segmentation and semantic verification. The method comprises the following steps: receiving a conference audio stream, identifying effective audio frames based on dual-channel voice activity detection, and pressing the effective audio frames into a dynamic buffer area; according to a preset dynamic segmentation strategy, outputting an initial voice segment from the dynamic buffer area for voice recognition, and obtaining a corresponding initial text segment; performing multi-stage reliability verification on the initial text fragment, and dynamically correcting or complementing the initial text fragment according to a verification result to obtain a reliable recognition result; and after the reliable identification result is obtained, triggering an asynchronous parallel translation task. According to the method, a semantic correction mechanism cooperating with the adaptive truncation depth is designed, so that the semantic fragmentation problem of the long-sequence audio during streaming truncation is solved, and the translation accuracy of a complex word order language is greatly improved while low delay is ensured.
Owner:WUHAN UNIV

Interactive labeling method for 3D dynamic object based on time series data, key frames, and interpolated frames

Disclosed is interactive labeling of a 4D dynamic object based on time series data, which aims at time series-related point cloud dynamic object data. Multi-frame local point clouds in the same time series are transformed into the same global coordinate system with corresponding poses to obtain global point clouds in the same time series are obtained, which clearly shows the moving trajectory of the dynamic object. Taggers can label key 3D boxes based on the moving trajectory of the dynamic object, and automatically generate 3D prediction boxes of other frames based on these key 3D boxes, which significantly reduces the number of frames that need manual operation and solves the problem that 3D prediction boxes generated based on deep learning model are inaccurate and efficiency can hardly be improved.
Owner:MOLAR INTELLIGENCE INFORMATION TECHNOLOGY (HANGZHOU) CO LTD

Residual frame video processing convolution method based on dynamic gating

The invention relates to the technical field of computer vision and video processing, and provides a residual frame video processing convolution method based on dynamic gating, which comprises the following steps of: carrying out gray conversion and Gaussian filtering processing on a current frame image and a previous frame image; calculating an inter-frame difference image and carrying out binarization processing to obtain a motion area mask; performing morphological expansion operation on the mask image; extracting the contour of the motion area and calculating a minimum bounding rectangle to obtain a bounding box; carrying out merging processing on the overlapped bounding boxes; updating the synthesized background frame based on the combined bounding box, and copying the pixels of the motion area of the current frame to the corresponding position of the background frame; and outputting the optimized bounding box set and the updated synthesized background frame. According to the method, the calculation efficiency and the resource utilization rate of video processing are improved, and meanwhile, the timing sequence continuity and the space consistency of a processing result are ensured through a strategy of keeping the static region unchanged and only updating the motion region.
Owner:CHANGSHA CHENGZHUO MICROELECTRONICS CO LTD

Information processing apparatus, information processing method, and computer-readable storage medium

The application discloses an information processing device and method based on zero-order learning and a computer readable storage medium. The information processing device comprises: an activation heat map generation unit configured to generate an activation heat map of each frame in a video based on a predetermined request by using a pre-trained multi-modal model or a pre-trained attention network; a spatial region of interest determination unit configured to determine a region of interest of each frame based on the activation heat map of the frame; and a similarity time sequence obtaining unit configured to calculate a similarity between the region of interest of each frame and the predetermined request by using the pre-trained multi-modal model to obtain a similarity time sequence of the video, which can be used to identify a target frame corresponding to the predetermined request in the video.
Owner:FUJITSU LTD

Local editing large model jailbreak attack method based on activation guidance

The invention relates to the technical field of large language models, and discloses a local editing large model jailbreak attack method based on activation guidance, which comprises the steps of data acquisition and marking work, data preprocessing and feature extraction, two-stage jailbreak attack based on an AGILE framework, a generation stage and an editing stage, and generation of a final jailbreak prompt. Malicious queries are converted into hidden jail break prompts based on a two-stage framework, and efficient jail break attacks are realized by guiding editing through activation signals in a model. According to the method, through AGILE two-stage framework design, in the generation stage, multiple rounds of dialogue history H are constructed at a time by means of a generator LLM, and in the editing stage, activation and attention scores are utilized to guide subtle and local editing of a generated text, so that two-stage decoupling design is achieved, the expandability of attacks is effectively improved, and the method has the advantages of being high in practicability and easy to popularize. In the generation stage, a generator LLM only needs to be called once, and efficient optimization is achieved by means of lightweight operation in the editing stage.
Owner:HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS +1

Correction method for virtual scene displacement missing for mixed reality

The invention provides a virtual scene displacement missing correction method for mixed reality. The method comprises the following steps: acquiring respective environment association information of a current frame and a previous frame of the current frame; determining whether a displacement correction operation for the current frame is triggered or not before the current frame is displayed based on the environment association information; when it is determined that the displacement correction operation for the current frame is triggered, obtaining a segmentation index of the current frame, and querying a displacement proportionality coefficient and a rotation proportionality coefficient corresponding to the segmentation index in a preset parameter statistics calibration table; determining a predicted displacement of the current frame based on the displacement proportionality coefficient, and determining a predicted rotation deviation of the current frame based on the rotation proportionality coefficient; determining an updated pose of the current frame based on the current frame, the previous frame, the predicted displacement and the predicted rotation deviation; and if the updated pose meets the preset correction amplitude limiting condition, performing displacement correction processing on the current frame based on the updated pose. According to the scheme, prediction and advanced compensation of a positioning mutation problem caused by environment mutation are realized.
Owner:BEIJING ZHIHUI HUANYU TECHNOLOGY CULTURE CO LTD

VVC reconstruction frame post-processing method based on large model

The invention relates to a VVC reconstructed frame post-processing method based on a large model, and belongs to the technical field of video coding and image processing. In order to solve the problems of compression artifacts and distortion existing in a VVC coding reconstruction frame, preliminary distortion suppression is performed through a loop filtering module, and coding side meta information is obtained; the input preprocessing module extracts visual features and meta-information features; the large model processing module fuses features by using a convolution attention mechanism, and realizes local enhancement and global modeling through a lightweight window self-attention and depth separable convolution hybrid network; and the output post-processing module adopts adaptive residual adjustment to generate a high-quality reconstructed frame. According to the method, the visual quality of the reconstructed frame can be remarkably improved, accurate self-adaptive repair is realized, good balance is achieved between high efficiency and light weight, and practical application deployment is facilitated.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Traditional art 3D content conversion method based on artificial intelligence

The invention discloses a traditional art 3D content conversion method based on artificial intelligence, and relates to the technical field of artificial intelligence, and the method comprises the following steps: constructing a cross-scene semantic anchor point graph, generating a verifiable semantic anchor point signature for the character features and background materials in a traditional art image, and carrying out the verification of the semantic anchor point signature; synchronously establishing a space-time constraint baseline based on space coordinates and a time sequence; and before scene switching, loading a millisecond preheating frame based on a space-time constraint baseline, executing time sequence calibration on the semantic anchor point signature, and outputting a phase reference for controlling a subsequent mapping process. According to the method, through semantic anchor point signature, space-time constraint, preheating frame calibration, hierarchical binding, mapping snapshot, texture springback, residual error remapping and a man-machine resonance regulation and control mechanism, precise binding and dynamic stability of semantic tags in a three-dimensional space are achieved, and the reduction degree of traditional art 3D conversion and the immersion continuity of multi-scene interaction are improved.
Owner:CHAOHU UNIV

Multi-target tracking method for self-adaptive threshold and buffer association

The invention provides a multi-target tracking method based on self-adaptive threshold and buffer correlation, which relates to the technical field of computer vision and comprises the following steps: calculating a self-adaptive confidence threshold of each frame of image, dividing a detection frame of each frame of image into a high-score detection frame and a low-score detection frame, generating an initial motion track set, and tracking the initial motion track set according to the self-adaptive confidence threshold; performing two-stage matching on the detection frame of each frame of image and the tracks in the initial motion track set, and generating a motion track set of each frame of image based on a matching result, the motion track set including a plurality of matched tracks; and generating a target tracking result of each frame of image based on the movement track set of each frame of image. According to the method, an adaptive confidence threshold mechanism and a buffer intersection-union matching strategy are introduced, so that the threshold can change in real time according to the scene, and limitation caused by a fixed threshold is avoided. And a buffer intersection-parallel ratio is introduced, the boundary of the low-score detection frame is properly expanded, and the expanded buffer frame is still possibly overlapped with the trajectory prediction frame sufficiently, so that successful matching is realized.
Owner:NORTHEASTERN UNIV CHINA

Target tracking method based on two-stage selection

The invention relates to a target tracking method based on two-stage selection, and the method comprises the following steps: obtaining an input video stream sequence, extracting image features for a current frame, and generating at least one candidate mask and a confidence score corresponding to each candidate mask based on memory bank information; predicting the motion state of the target in the current frame based on a motion prediction model, and calculating a coincidence evaluation score between each candidate mask and the predicted motion state; selecting a first candidate result from the candidate masks based on the confidence score and the anastomosis evaluation score, and judging whether the anastomosis evaluation score of the first candidate result meets a preset stability condition or not; and if yes, determining the first candidate result as a final segmentation result of the current frame, and if not, calculating a similarity evaluation score between each candidate mask and a historical segmentation result stored in a historical state library, and selecting a second candidate result from the candidate masks as the final segmentation result of the current frame based on the similarity evaluation score.
Owner:TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL +1

A method and system for determining whether a beaker experiment is scored

The application provides a method and system for judging whether a beaker experiment is successful, comprising: obtaining a video stream of the beaker experiment and performing region of interest detection; a plurality of images containing the region of interest obtained by the region of interest detection constitute a picture set of a region of interest period of the beaker experiment; a full convolutional neural network (FCN) model is trained based on the picture set of the region of interest period of the beaker experiment; the picture set of the region of interest period is subjected to strong supervision target classification frame by frame based on the trained FCN model, and a classification result of each frame in the picture set of the region of interest period is obtained; and whether the experiment operation in the video stream is successful is determined based on the classification result of each frame in the picture set of the region of interest period. The data required by the application is greatly reduced, and the application is suitable for video data size in actual experiment operation process; and the target region can be quickly focused on by detecting the region of interest, and the detection precision is effectively improved.
Owner:SHANGHAI MEDIA INTELLIGENCE TECH CO LTD

Artificial intelligence system for big data processing

The invention belongs to the field of data processing, and discloses an artificial intelligence system for big data processing, which comprises an initial detection module, a motion prediction module, a detection frame generation module and a target detection module. The initial detection module is used for detecting a moving target based on all pixel points in the first two frames in the frame sequence to obtain state information of the moving target in the second frame; the motion prediction module is used for predicting the state information of the moving target in the nth frame based on a Kalman filtering algorithm and the state information of the moving target in the (n-1) th frame, n belongs to [3, N], and N represents the total number of frames in the frame sequence; the detection frame generation module is used for generating a detection frame of the nth frame based on the state information; and the target detection module is used for performing moving target detection on the nth frame based on the detection frame to obtain an area where the moving target is located. According to the invention, the accuracy of the detection result can be ensured while the calculated amount of moving target detection is effectively reduced.
Owner:GUANGDONG ACADEMY OF SCIENCES SHANWEI IND TECHNOLOGY RESEARCH INSTITUTE CO LTD

Spring framework-based user behavior tracing and operation log auditing method and platform

The invention provides a Spring framework-based user behavior tracing and operation log auditing method and platform. The method comprises the following steps of: obtaining a user operation request; analyzing the user operation request to obtain an event sequence; on the basis of the event sequence, adjacent events are connected according to a time sequence to generate directed edges, the directed edges are organized into a semantic path in combination with the event sequence, and a risk weight is given to cross-module jump operation; according to the semantic path, combining a preset legal operation path rule set to design a path matching scoring function to perform legality verification, and outputting a path legality identifier and a structural abnormal position set in an illegal path; and generating an operation log in combination with the semantic path, the path legality identifier, the structural abnormal position set in the illegal path and the abnormal context mapping function. According to the method, the problems of behavior chain missing, judgment lagging and insufficient control capability in a traditional system are practically solved.
Owner:GUANGZHOU DUOMI TECHNOLOGY DEVELOPMENT CO LTD

Method and device for recognizing an object, as well as machine learning model and computer program product

PendingDE102024133111A1Character and pattern recognitionFrame basedObject type
The present disclosure relates to a method for recognizing an object. The method comprises the following steps: i) Comparing a detected object from a second frame with at least one detected object from a first frame using at least one of the following parameters: a) distance in terms of distance, in particular the Euclidean distance between the objects; b) Difference between the objects in their geometric dimensions, in particular the height and / or width of the objects; c) Difference between the objects in their preferably optical and / or acoustic appearance; d) Equality of the objects with respect to an object type; ii) Evaluating the results from step i) to determine whether the detected object of the second frame is the detected object of the first frame. The present disclosure further relates to a device (1) for recognizing an object, as well as a machine learning model and a computer program product.
Owner:BEEBUCKET GMBH

Systems and methods for anti-aliasing

ActiveUS12555282B2Image enhancementDetails involving antialiasingPattern recognitionFrame based
There is provided a computer-implemented method for anti-aliasing an image frame in an image stream. The method comprises providing an input to a data model, the data model being configured to perform anti-aliasing on a present frame which is displayed after a previous frame in the image stream. The input comprises: an anti-aliased image of the previous frame; an aliased image of the present frame; an ID map of the previous frame; an ID map of the present frame; and a velocity map based on changes between the previous frame and the present frame. The data model processes the received inputs to correlate image portions of the previous frame and image portions of the present frame based on the inputs and performing anti-aliasing on the present frame based at least in part on the correlated image portions. An anti-aliased image of the present frame is received as an output of the data model.
Owner:SONY COMP ENTERTAINMENT EURO LTD

Computer vision-based automatic ore-drawing control method and system

The application provides a computer vision-based automatic ore-drawing control method and system, which collects real-time images of a drawing bin, identifies whether a target vehicle exists in the real-time images, generates a detection frame based on the target vehicle in the case where the target vehicle exists in the real-time images, and determines whether a center point of the detection frame is located in a center region in the case where the target vehicle is in a stopped state; wherein the center region is a preset region in the real-time images, and ore is drawn to the target vehicle by a ore-drawing machine in the case where the center point is located in the center region, so that the artificial cost can be reduced, the mineral loss in the ore-drawing process can be reduced, and the ore-drawing efficiency can be improved.
Owner:CHENGDU SUNLIGHT TECH CO LTD

A double-flow end-to-end visual autonomous positioning method and system based on bio-inspired spatio-temporal reasoning

PendingCN122089835AEfficient long-range motion modelingImage analysisBiological modelsFeature vectorVisual technology
This invention relates to a two-stream end-to-end visual autonomous localization method and system based on bio-inspired spatiotemporal reasoning, belonging to the field of computer vision technology. The method includes: acquiring an input image sequence containing adjacent frames; inputting the input image sequence into a dorsal visual pathway to extract and aggregate inter-frame motion features to obtain a high-dimensional motion feature vector; inputting the input image sequence into a ventral visual pathway to extract inter-frame spatial correlation features through adaptive temporal convolution and a visual state space module to obtain spatial semantic features; performing feature fusion and interaction through a gated cross-attention fusion module, using spatial semantic features as queries and high-dimensional motion feature vectors as keys and values; and estimating the relative pose between adjacent frames based on the fused features. This scheme achieves accurate and robust pose estimation without requiring camera intrinsics, striking a balance between inference speed, model complexity, and training cost, making it suitable for deployment in practical robot applications.
Owner:CHONGQING UNIV +1

A computer organization and principle course auxiliary question and answer method, device and medium

This invention discloses a question-answering method, device, and medium for a computer organization principles course, relating to the field of natural language processing technology. The method includes: First, integrating multi-source data, cleaning and format conversion to construct a question-answering dataset for training. Next, supervising fine-tuning of the ChatGLM3-6B model based on the LoRa method, while simultaneously vectorizing course documents using a semantic vector model to construct a text vector knowledge base containing a vector database and a custom code library, supporting document updates and code retrieval. Finally, a WebUI interface is built based on the LangChain-Chatchat framework, integrating the fine-tuned large model and vector knowledge base, providing users with two question-answering modes. This invention enables efficient fine-tuning of the model based on domain knowledge and constructs a corresponding question-answering assistance mechanism to provide professional, accurate, and tailored intelligent support for teaching needs.
Owner:YANSHAN UNIV

Augmented reality navigation image rendering method and device, electronic device, and storage medium

The application discloses an AR navigation image rendering method and device, electronic equipment and a storage medium, and is used for rendering a real scene image in a video frame. The method comprises the following steps: acquiring a current frame in the video frame; rendering the current frame, and recording a starting time and an ending time of the rendering; acquiring first driving state information and second driving state information of the vehicle at the starting time and the ending time respectively; calculating the latitude and longitude of the vehicle when the next frame is acquired based on the starting time, the first driving state information, the ending time, the second driving state information and the latitude and longitude of the vehicle when the current frame is acquired; and rendering the next frame based on the calculated latitude and longitude of the vehicle when the next frame is acquired. The application reduces the algorithm complexity and the power consumption, reduces the calculation error and improves the accuracy. The latitude and longitude are used as the reference to participate in the modeling calculation, so that the model can be real-time fitted with the actual situation.
Owner:SHANGHAI PATEO INTERNET TECH SERVICE CO LTD

Driving mode judgment method and system based on computer vision

The invention discloses a driving mode judgment method and system based on computer vision, and relates to the technical field of driving mode judgment, and the method comprises the following steps: extracting a static image frame, and dynamically adjusting the frame extraction frequency; carrying out local contrast self-adaptive adjustment on the image frame to obtain an enhanced image; extracting and fusing a brightness gradient and texture distribution based on multi-level visual feature analysis to generate a unified feature vector; inputting the unified feature vector into an intelligent discrimination model for navigation picture recognition and outputting a result with confidence; and carrying out dynamic weighted updating on the confidence of the continuous frames based on time sequence attenuation and a sliding window mechanism. Through computer vision identification and time sequence regulation and control, intelligent identification and dynamic release of the navigation picture by the vehicle-mounted terminal are realized, interruption caused by error limit is avoided, navigation continuity and driving safety are guaranteed, and identification accuracy and response sensitivity are kept in stable and sudden change scenes through an adaptive adjustment mechanism, so that the navigation effect is improved. And the intelligence and user experience of the vehicle-mounted system are improved.
Owner:SHANGHAI JIDOU TECH CO LTD

Video frame extraction method and device, equipment, storage medium and program product

The invention provides a video frame extraction method and device, equipment, a storage medium and a program product, and relates to the technical field of artificial intelligence. The method comprises the following steps: detecting a target abrupt change frame according to the color difference degree and the feature point distance of image blocks between frames in a video frame sequence; detecting a target gradient frame group according to the brightness change between the frames and the matching number of the feature points; according to the target mutation frame and the target gradual change frame group, segmenting the video frame sequence to obtain a plurality of sub video frame sequences; and extracting a target video frame based on the inter-frame difference intensity in the sub-video frame sequence. According to the method, the accuracy of extracting the video frame is improved.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Methods, apparatus, equipment and storage media for adjusting the number of feature points

This application discloses a method, apparatus, device, and storage medium for adjusting the number of feature points, relating to the field of intelligent device technology. The method includes: when the timestamp of the previous frame construction is detected to be initialized, calculating the total time taken to construct the previous frame based on the timestamp; calculating the total time taken to construct the current frame based on the timestamp; calculating a target total time based on the total time taken to construct the previous frame and the total time taken to construct the current frame; when the target total time is greater than a preset single-frame period threshold, adjusting the number of feature points in the next frame by combining the number of feature points in the original frame construction and a preset adjustment coefficient. Through the above method, it detects whether the timestamp of the previous frame construction has been initialized. If so, the number of feature points is dynamically adjusted based on the comparison result between the target total time and the preset single-frame period threshold. This method is applicable to various complex scenarios, ensuring the stability and real-time performance of pose output.
Owner:GOERTEK INC

Pre-training language model post-interpretation method and system based on framework knowledge detection

PendingCN121543591ASemantic analysisKnowledge representationFrame (artificial intelligence)Frame based
The invention discloses a pre-training language model post-interpretation method and system based on framework knowledge detection, and particularly belongs to the field of artificial intelligence and natural language processing. Comprising a framework-based semantic analysis module, a framework-based knowledge graph construction module and an interpretable knowledge detection module. The method is a post-interpretation mechanism based on a framework semantic linguistics cognitive mechanism, and the core of the method is to evaluate and interpret internal parameter concepts of a pre-training language model which cannot be understood by human beings by using framework cognitive concepts which can be understood by the human beings. By introducing a framework, framework elements and framework relationship concepts, a framework knowledge graph capable of representing implicit knowledge behind a text is realized based on a given text. And then a knowledge detection module is combined to generate multiple types of detection problems, and understanding and reasoning of the model on text back implicit knowledge are evaluated. According to the method, the limitation of the model in the aspects of knowledge representation and reasoning can be effectively identified, and support is provided for model optimization and research of a credible artificial intelligence model.
Owner:SHANXI UNIV

AI responsive layout for cross-platform environments

A method includes executing a video game to generate video frames for presentation on a device of a first platform for a user to play the game. The method includes determining a target device of a second platform. The method includes mapping a video frame to a target device. The method includes determining a focus region in a scene in a video frame based on a game context. The method includes classifying assets in a focal region using a computer vision model that implements artificial intelligence. The method includes using a computer vision model to determine that the asset is important during gameplay. The method includes determining that an asset in the mapped video frame does not satisfy a visibility threshold. The method includes modifying the mapped video frame such that the asset meets a visibility threshold.
Owner:SONY INTERACTIVE ENTERTAINMENT LLC

Method and apparatus for automatically aligning cad bounding boxes with thumbnails

The application provides a CAD detection frame and thumbnail automatic alignment correction method and device, generates a geometric standard image matched with a CAD coordinate system, identifies a plurality of mark points preset on the thumbnail in the geometric standard image, filters out an effective reference point set satisfying a registration accuracy threshold value by calculating the spatial Euclidean distance of the mark points and CAD theoretical mark points, constructs a similarity transformation model of the coordinate system according to the image coordinates and CAD theoretical coordinates of the effective reference point set, determines the translation, rotation and scale constraint parameters, performs multi-parameter coupling iterative correction on the pose of the detection frame based on the pose deviation characteristics of the detection frame in the geometric standard image and the similarity transformation model parameters, obtains the pose correction parameters of the detection frame, updates the pose optimization parameters to the detection frame attributes, and realizes the automatic alignment of the detection frame and the thumbnail. The scheme of the application can realize multi-parameter coupling alignment correction of the CAD detection frame and the thumbnail.
Owner:SHENZHEN ZHENHUAXING INTELLIGENT TECH CO LTD

Region-based filtering

Various embodiments describe apparatus, method and computer program product. An example apparatus includes at least one processor; and at least one non-transitory memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to perform: receiving a frame to be filtered; receiving semantic information about the frame to be filtered; and filtering the frame based at least on the semantic information.
Owner:NOKIA TECHNOLOGIES OY

A data labeling method, device, apparatus, and readable storage medium

The application discloses a data labeling method and device, equipment and a readable storage medium, which can be applied to the field of artificial intelligence technology, and automatically labels the composition frame score of each candidate frame based on a pre-trained aesthetic large model, performs first labeling to obtain a high-score composition image, performs second labeling based on the sample labeling score of the high-score composition image fed back by a client, obtains multiple optimal composition frames and corresponding labeling score results, the first labeling is realized based on automatic labeling, and the second optimal selection is realized based on a small amount of manual labeling, that is, it is not necessary to score each candidate frame manually, the pre-trained aesthetic large model and a small amount of manual intervention are used to obtain a labeling training set for retraining the aesthetic large model, the aesthetic large model is fine-tuned based on the labeling training set, the artificial cost of constructing the aesthetic large model is reduced, and the accuracy of automatic labeling of the aesthetic large model is improved.
Owner:SHENZHEN LINKRIC TECH CO LTD