Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

133 results about "Scene segmentation" patented technology

Remote sensing scene graph guided semantic information reasoning method and device, equipment and medium

The invention provides a semantic information reasoning method guided by a remote sensing scene graph, which can be applied to the technical field of remote sensing image processing. The method comprises the following steps: segmenting a remote sensing image to generate a scene segmentation image; executing target detection to generate an image block set; performing fine-grained analysis on the image global features and the image block set to generate a scene description text; performing grammar analysis, extracting a triple of objects, object attributes and relation information among the objects, and generating a remote sensing scene graph related to the problem; encoding the problem text into an embedded vector, projecting the embedded vector to a visual feature space, processing the embedded vector through a frozen self-attention layer and a learnable gating layer, and outputting an image-text interaction feature; carrying out attention fusion and double gating balance on the question coding features, and outputting a text code guided by a scene graph; and fusing the interaction features and the text codes, and reasoning local semantic information of the target area. The invention further provides a semantic information reasoning device and equipment guided by the remote sensing scene graph and a medium.
Owner:AEROSPACE INFORMATION RES INST CAS

Building construction intelligent safety monitoring method based on Internet of Things

The invention discloses a building construction intelligent safety monitoring method based on the Internet of Things, and the method comprises the following steps: obtaining construction environment data, structure state data and operation behavior image data, and carrying out the preprocessing; environment state modeling, local structure strain and behavior recognition and operation scene segmentation are carried out through the edge intelligent processing unit; feature fusion is carried out, and a fusion situation vector is constructed; performing high-frequency anomaly identification and emergency preliminary screening, and generating an edge preliminary early warning result and a high-risk data fragment; constructing a safety evolution trajectory crossing a time window, and fusing historical data to generate a risk semantic map; cloud semantic reasoning operation is executed, and a final risk level judgment result and a corresponding trigger source identifier are generated; and automatically triggering a safety response instruction according to a risk level judgment result. According to the method, the Internet of Things and intelligent semantic analysis are fused, multi-source safety monitoring and self-adaptive response are realized, and the method has the advantages of high real-time performance, global perception and continuous optimization.
Owner:GUIZHOU CONSTRUCTION GROUP CHONGQING GUIYU CONSTRUCTION CO LTD

Substation inspection robot multi-sensor fusion obstacle avoidance method and system

The invention belongs to the technical field of substation inspection, and provides a substation inspection robot multi-sensor fusion obstacle avoidance method and system, and the technical scheme is as follows: obtaining multi-modal sensor data, including laser point cloud data, millimeter wave radar point cloud data, ultrasonic radar observation data and visual image data; preprocessing the obtained multi-modal sensor data to obtain effective observation data, and determining initial obstacle data based on the effective observation data; time synchronization and fusion are carried out on the obtained multiple kinds of initial obstacle data, the fused data serve as obstacle data, scene segmentation is carried out in combination with multiple kinds of sensor data in the fusion process, confidence coefficient decision rules are dynamically set according to different scenes, and obstacle information is determined in combination with the set confidence coefficient decision rules; and generating an optimal obstacle avoidance route based on the determined obstacle information. The problem that a traditional single sensor cannot cope with a complex scene of substation inspection business is solved, and the overall stable operation capacity of the robot is improved.
Owner:STATE GRID INTELLIGENCE TECHNOLOGY CO LTD

Double-arm clothes folding robot control method, device and equipment and storage medium

The invention discloses a double-arm clothes folding robot control method and device, equipment and a storage medium. The method comprises the following steps: determining three-dimensional point cloud data of a target workbench according to a collected image through a visual perception assembly; determining a scene label graph and a desktop edge line set through a scene division assembly according to the three-dimensional point cloud data, determining a clothing mask of the to-be-processed clothing according to the scene label graph through a clothing classification assembly, and determining a clothing description vector according to the clothing mask; determining a low-risk action set and a protection area mask from the candidate clothes folding action set through a safety protection component according to the scene label graph, the desktop edge line set and the clothes mask; determining a target clothes folding action sequence through an action planning component according to the clothes description vector, the low-risk action set and the protection area mask; and folding the to-be-processed clothes through the robot control assembly according to the target clothes folding action sequence. According to the scheme, the clothes folding efficiency and reliability of the robot are improved.
Owner:KAILONG HIGH TECH CO LTD +2

Yolo-based night smoke and fire identification optimization method

The invention belongs to the technical field of computer vision, embedded edge calculation and intelligent video monitoring, and discloses a yolk-based night smoke and fire identification optimization method, which can effectively detect flames at night or in a scene with disordered light. A multi-source data fusion and lightweight scene segmentation technology is adopted, an on-site perception model is built, and the model is combined with a small sample learning mechanism, so that night common interference factors such as vehicle lamps and street lamps can be identified; then, a federal learning algorithm and a time sequence attention module are used for modeling the dynamic characteristics of the smoke and fire, and an obtained characteristic model has high generalization ability; then, the system accurately screens out real smoke and fire candidate targets through self-adaptive threshold matching and pixel-level interference elimination operation; and finally, based on a multi-factor decision model and a dynamic edge transmission technology, rapid alarm of smoke and fire identification is realized. The whole system can also automatically optimize the whole process parameters of the method through rule mining and reinforcement learning.
Owner:SUZHOU BIANCHI TECHNOLOGY CO LTD

Data security monitoring and analysis method for bidding informatization platform

The invention relates to the technical field of information monitoring, in particular to a bidding informatization platform data security monitoring and analysis method, which comprises the following steps of: acquiring bidding platform data, dividing user behaviors into a plurality of monitoring points according to operation objects, operation quantity and operation types of the bidding platform data in a continuous time period, acquiring a user behavior sequence formed by the monitoring points; performing scene segmentation on the user behavior, and determining an abnormal behavior relative to a historical behavior baseline under each scene segmentation dimension; merging each abnormal behavior into a plurality of association sets, and determining a behavior association result of each association set and the operation object; determining association paths under different scene segmentation dimensions according to a behavior association result in a monitoring point distribution form; for the intersection of the association paths under all scene segmentation dimensions, updating a path list after combination of the association paths, and taking the updated path list as an output early warning target; the response speed and the processing efficiency of the platform to the abnormal behavior are realized.
Owner:CHINA COAL INFORMATION TECH (BEIJING) CO LTD

Hidden space world model construction method for robot grabbing operation scene and related equipment thereof

The invention belongs to the technical field of artificial intelligence, and discloses a hidden space world model construction method for a robot grabbing operation scene and related equipment thereof. The method comprises the steps of performing semantic segmentation on a scene image through an image semantic segmentation network to generate a semantic segmentation mask of each object; processing the semantic segmentation mask through a multi-task scene understanding network to generate implicit representation; based on the implicit representation of each object at the current moment and the action information of the mechanical arm, predicting implicit representation at the next moment through a state transition network; and according to the predicted implicit representation at the next moment, generating a scene segmentation image reconstruction result, an object existence judgment result and a contact relation judgment result through a decoder in the multi-task scene understanding network architecture, and constructing an implicit space world model of the scene. Based on the method, efficient characterization and dynamic prediction of the robot grabbing operation scene are achieved, and the generalization ability and the multi-task cooperation efficiency of the model in a complex environment are improved.
Owner:QINGDAO UNIV OF TECH

Self-adaptive sensing strategy switching method and system for long-distance automatic driving

The invention relates to a self-adaptive perception strategy switching method and system for long-distance automatic driving, and the method comprises the steps: carrying out the analysis of a public data set through employing a multi-target joint optimization clustering method, generating a standard scene prototype, and constructing an initial scene-optimal strategy mapping database; acquiring multi-dimensional scene features of a navigation route, performing global optimal scene division and recognition on a target road section, and generating a perception strategy map between continuous cells; querying a scene-optimal strategy mapping database continuously optimized by reinforcement learning, and selecting an optimal perception strategy for each cell based on a Q value maximization principle; when the vehicle enters the new cell, the self-adaptive switching control module automatically activates the preloaded optimal sensing strategy; and evaluating a strategy execution effect, generating an instant reward, transmitting an experience tuple containing a state, an action and the reward to the cloud, and continuously optimizing the Q-value database. The method and the system can improve the perception performance and the resource utilization efficiency of the automatic driving system.
Owner:FUJIAN NORMAL UNIV

Deep neural network for segmentation of road scenes and animate object instances for autonomous driving applications

A deep neural network(s) (DNN) may be used to perform panoptic segmentation by performing pixel-level class and instance segmentation of a scene using a single pass of the DNN. Generally, one or more images and / or other sensor data may be stitched together, stacked, and / or combined, and fed into a DNN that includes a common trunk and several heads that predict different outputs. The DNN may include a class confidence head that predicts a confidence map representing pixels that belong to particular classes, an instance regression head that predicts object instance data for detected objects, an instance clustering head that predicts a confidence map of pixels that belong to particular instances, and / or a depth head that predicts range values. These outputs may be decoded to identify bounding shapes, class labels, instance labels, and / or range values for detected objects, and used to enable safe path planning and control of an autonomous vehicle.
Owner:NVIDIA CORP

Scene segmentation and editing method and system based on three-dimensional Gaussian sputtering

The invention provides a scene segmentation and editing method and system based on three-dimensional Gaussian sputtering, and the method comprises the steps: carrying out the two-dimensional instance segmentation of a three-dimensional scene multi-view two-dimensional image, obtaining a two-dimensional instance mask, initializing the three-dimensional reconstruction through three-dimensional Gaussian sputtering, and obtaining a three-dimensional Gaussian body set; associating cross-view consistent semantic features extracted by the pre-training visual language model for each Gaussian body, constructing and iterating a three-dimensional instance prototype library based on a two-dimensional instance mask, and performing conversion to obtain a cross-view consistent two-dimensional supervision signal; adding a learnable instance identity code for each Gaussian body, constructing a multi-dimensional loss function in combination with a two-dimensional supervision signal, a spatial neighborhood relationship and cross-view consistent semantic features, and optimizing learnable parameters containing the instance identity codes; and realizing Gaussian body instance grouping based on the optimized instance identity codes to complete instance-level scene segmentation and editing. According to the method, the technical problems of instance identity association and high-quality instance level generation under the conditions of sparse view angle and close adjacent objects are solved.
Owner:NANCHANG CAMPUS OF JIANGXI UNIV OF SCI & TECH

Laparoscopic surgery full-granularity recognition system and method based on feature extraction and task segmentation

The invention belongs to the technical field of medical information processing, and relates to a laparoscopic surgery full-granularity recognition system and method based on feature extraction and task segmentation. Firstly, a video stream composed of multiple frames of images of an endoscope is input into the system, and then the spatial feature of each frame of image is obtained through a multi-frame spatial feature extraction module; the spatial features of all the frames are then input into a multi-frame time feature fusion module to obtain time features fused with different time lengths, and then the time features are used for outputting operation stages, operation steps and operation ternary body labels; meanwhile, only the spatial features of the current frame are extracted from the spatial features of all the frames, and an operation scene segmentation label is output through a single-frame feature restoration module. According to the laparoscopic surgery full-granularity recognition system and method based on feature extraction and task segmentation, computing resources can be greatly reduced, and compared with an existing single-granularity surgery scene technology, the laparoscopic surgery full-granularity recognition system and method based on feature extraction and task segmentation can output surgery scene information of multiple granularities at a time by means of the mutual coordination effect of multiple tasks.
Owner:THE AFFILIATED HOSPITAL OF QINGDAO UNIV

Method and processing device for providing a scene segmentation map

A method and processing device for providing a scene segmentation map for a total field of view of an image capturing device comprises obtaining a scene segmentation map for the total field of view; receiving an indication that a new focus value has been set for the image capturing device for acquiring images of a current field of view; comparing the new focus value with a stored focus value associated with a focus region in the current field of view, wherein the stored focus value represents a previously used focus value; upon the new focus value deviating from the stored focus value more than a trigger threshold, triggering a scene segmentation map update process; and otherwise maintaining the scene segmentation map.
Owner:AXIS

Ultrahigh-resolution immersive scene rendering method based on hexagonal picture segmentation

The invention discloses an ultrahigh-resolution immersive scene rendering method based on hexagonal picture segmentation. The method comprises the following steps: 1, determining a hexagonal segmentation style of an immersive display surface; 2, unfolding the immersive display surface into a plane; 3, acquiring 3D point information of each hexagonal surface, respectively calculating the width W and the height H of the enveloping quadrilateral surface, and calculating a local coordinate system V of the hexagonal surface S; 4, using W, H and V to construct a view cone and a viewpoint matrix of the three-dimensional scene; 5, setting a shade M of a plane expansion view for the hexagonal surface S; 6, calling a rendering scene of the 3D scene for the mask region by using information of the view cone and the viewpoint matrix, and rendering the mask region into a picture at the position S of the hexagonal surface; and 7, repeating the steps 5-6, and sequentially carrying out the same treatment on each hexagonal surface. According to the method and the device, the ultrahigh-resolution picture can be rendered in a scene segmentation form, and the method and the device are suitable for ultrahigh-resolution picture rendering under immersive display equipment.
Owner:HARBIN AIWELL TECH CO LTD

Indoor point cloud scene open type semantic segmentation method based on text prompt

The invention relates to an indoor point cloud scene open type semantic segmentation method based on text prompt. On the basis of large-scale indoor scene point cloud data, a 3D U-Net network structure based on Mama blocks is adopted to extract local features of a point cloud scene, global features of the scene point cloud data are extracted through multi-stage Mama blocks and down-sampling, and scene detail information is recovered through up-sampling and jump connection to fuse the local-global features of the point cloud data; generating a text title prompt of a scene multi-view view, associating a projection matrix between a 2D view and a 3D scene with a point cloud scene, enabling text prompt features to be aligned with corresponding point cloud data features, and loading a category text embedded weight into a scene segmentation head to perform a semantic segmentation task; in network training, binary classification loss is added on the basis of semantic segmentation loss so as to balance the recognition and understanding capability of the scene semantic segmentation network on a basic category and a new category; and finally, open semantic segmentation of the indoor point cloud scene is realized.
Owner:HANGZHOU NORMAL UNIVERSITY

Automatic generation method and device of plot abstract and electronic equipment

The invention relates to an automatic generation method and device of a plot abstract and electronic equipment. The method comprises the following steps: performing scene segmentation and key plot anchor point identification on each single-set original script of a target video, and generating a corresponding single-set draft abstract by using a large language model; performing multi-dimensional quality evaluation and screening processing on each single-set draft abstract, and converting the single-set draft abstract into a structured data object; sequentially arranging all the data objects to obtain a single-set abstract sequence of the target video; inputting the single-set abstract sequence into a large language model to enable the large language model to generate a whole draft abstract of the target video in combination with a preset narrative structure guide prompt; and performing multi-dimensional quality evaluation on the whole drama draft abstract, and optimizing the whole drama draft abstract based on an evaluation result to obtain a drama abstract of the target video. According to the method and the device, the technical problem that the effect and the reliability of an existing method in automatic abstracting of the long series are poor is solved.
Owner:BEIJING QIYI CENTURY SCI & TECH CO LTD

Marine scene segmentation method based on condition invariant semantic segmentation

The invention discloses an offshore scene segmentation method based on condition invariant semantic segmentation, and belongs to the field of computer vision and semantic segmentation. According to the method, style mutual conversion of a source domain image and a target domain image is realized through a data loading stage, and a multi-style view pair is generated; a CISS network model is constructed, the CISS network model is composed of a MiT-B5 encoder and a context sensing decoder, the encoder extracts high-level features insensitive to visual conditions by means of discrete wavelet transform, and the decoder achieves feature reconstruction without information loss through inverse discrete wavelet transform and generates a category probability graph; designing a joint loss function including pixel-level cross entropy loss and feature-level invariance loss, and balancing loss weight in combination with a dynamic weighting strategy; and training and optimizing the model by adopting an AdamW optimizer in combination with a linear learning rate warm-up strategy. Experiments show that the method can effectively cope with visual condition changes in a real marine environment, and the robustness of the semantic segmentation model in different marine environments is improved.
Owner:STATE POWER INVESTMENT CORP JIANGSU OFFSHORE WIND POWER +1

Process planning model pre-training method and system based on multi-processing scene segmented imitation learning

The invention discloses a process planning model pre-training method and system based on multi-processing scene segmented imitation learning, and relates to the field of computer-aided manufacturing, comprising the steps of: for various processing scenes comprising different geometric features, material characteristics and processing requirements, generating processing paths of various strategies by using a mature process database of CAM software; virtual trial cutting is carried out through a processing process world model to generate a large number of initial states, rich and diversified training samples are provided for reinforcement learning pre-training, and on the basis, large-scale pre-training is carried out by adopting a segmented behavior cloning method to obtain a process autonomous planning base agent which learns processing strategies in different processing scenes; in the face of a new processing task, a processing path can be quickly generated through transfer learning. According to the method, the generalization ability is remarkably improved, meanwhile, geometric-physical-control collaborative planning is achieved, and an intelligent solution is provided for efficient planning of the modern manufacturing industry technology.
Owner:SHANGHAI JIAOTONG UNIV

Scene segmentation method and system based on multi-modal small sample learning

The invention discloses a scene segmentation method and system based on multi-modal small sample learning, and relates to the technical field of computer vision. The method comprises the following steps: acquiring an automatic driving video stream, and extracting image frames in the video stream; training and testing a scene segmentation model by using the data set to obtain a trained scene segmentation model, performing semantic feature extraction on the data set by using a lightweight backbone network, and then performing double-axis self-attention feature extraction and information correction processing on the image frame after semantic feature extraction by using the scene segmentation model to obtain a trained scene segmentation model; and finally, fusing the corrected multi-source features to obtain a prediction output graph. And performing scene segmentation processing on a real-time video stream to be segmented by using the trained scene segmentation model. According to the method, the understanding capability of semantic information in a complex scene can be effectively enhanced while the lightweight of the model is guaranteed, and the segmentation precision and generalization capability under the small sample condition are remarkably improved.
Owner:HEILONGJIANG UNIVERSITY OF SCIENCE AND TECHNOLOGY

Power transmission line detection method based on multi-modal data fusion and related equipment

The invention provides a multi-modal data fusion-based power transmission line detection method and related equipment. The method comprises the following steps: performing scene segmentation on a space region of a power transmission line based on geographic information and remote sensing image data to obtain a plurality of scene regions with space continuity; obtaining original multi-modal data of the power transmission line; screening from the original multi-modal data based on a preset distance condition to obtain target multi-modal data of the target area in the target time period; obtaining a graph structure representing a relationship between the target multi-modal data based on the target multi-modal data and the corresponding position information; performing feature extraction on the target multi-modal data to obtain target multi-modal features; performing feature fusion on the target multi-modal features based on the graph structure to obtain global fusion features; and determining the operation state of the target position of the power transmission line at the target moment based on the global fusion feature. The operation state of the power transmission line can be accurately and reliably determined.
Owner:FIBRLINK NETWORKS +4

Semi-supervised adaptive unstructured scene segmentation method

The invention discloses a semi-supervised adaptive unstructured scene segmentation method, and relates to the technical field of computer vision. The method comprises the following steps: acquiring label-free image data of an unstructured scene; inputting the weakly enhanced image into a teacher model to generate a first prediction result; inputting the strong enhancement image into the student model to generate a second prediction result; calculating the JS divergence between the first prediction result and the second prediction result, and dynamically adjusting the attenuation coefficient of the index moving average according to the JS divergence; updating teacher model parameters according to the adjusted index moving average attenuation coefficient to obtain an unstructured scene segmentation model; and inputting a to-be-predicted unstructured scene image into the trained unstructured scene segmentation model, and outputting an unstructured scene segmentation result. According to the method, the teacher model weight updating rate is reduced in the model divergence significant stage such as the initial training stage or the difficult sample processing stage so as to suppress noise propagation, and the prediction precision of unstructured scene segmentation is effectively improved.
Owner:JIANGSU XCMG STATE KEY LAB TECH CO LTD

A 3D Gaussian-based three-dimensional scene segmentation and interaction method

PendingCN122289679AImprove ability to respond accuratelyaccurate segmentationPattern recognitionGauss point
This invention discloses a 3D scene segmentation and interaction method based on 3D Gaussian, comprising: Step 1, instance discovery: inputting 3D Gaussian scene data, dividing the 3D Gaussian points into structurally coherent instance-level Gaussian groups to obtain refined instances; Step 2, instance and scene semantic assignment: selecting representative viewpoint images of refined instances and inputting them into a visual-language model to generate semantic description labels for the instances; filtering instance pairs in the scene through spatial geometric relationships, and clarifying the spatial relationship description of instance pairs through a large model, constructing a static scene graph that integrates geometric proximity relationships and instance semantic relationships, forming a structured description of all instances; Step 3, natural language-driven instance localization and interaction: receiving user input commands, combining the instance semantic description, the static scene graph, and the real-time viewpoint direction relationship during the query to perform multi-dimensional matching, determining the target instance ID, and executing interactive operations.
Owner:NANJING UNIV

Bicycle accounting evaluation system and method and electronic equipment

The invention provides a bicycle accounting evaluation system and method and electronic equipment, and relates to the technical field of bicycle data management, and the system comprises a data collection module which is used for obtaining the operation data flow of a bicycle in a preset time period; the scene classification module is used for carrying out dynamic clustering on the operation data flow of each bicycle through a scene segmentation model, generating a plurality of scene labels and obtaining a scene label set of the bicycles according to the scene labels; the loss mapping module is used for performing data mapping according to the scene label set and a preset loss cause library to obtain a loss matrix; the weight adjustment module is used for performing dynamic weight adjustment on the initial depreciation rate corresponding to the bicycle according to the loss matrix to obtain a dynamic depreciation weight vector; and the accounting generation module is used for carrying out apportionment calculation on a plurality of preset indexes of the bicycles according to the dynamic depreciation weight vectors, and generating an accounting result of each bicycle in the target area. According to the invention, the accuracy of bicycle accounting is improved.
Owner:STATE GRID ZHEJIANG ELECTRIC POWER CO LTD NINGBO POWER SUPPLY CO +2

Video scene segmentation method, apparatus, device, and storage medium

Embodiments of the present application disclose a video scene segmentation method, device, equipment and storage medium. The method comprises: performing video segmentation on the video to obtain a plurality of shot pictures; performing feature extraction on the plurality of shot pictures respectively to obtain multi-modal features of each shot picture in the plurality of shot pictures; performing fusion processing on the multi-modal features of each shot picture in the plurality of shot pictures respectively to obtain fusion semantic information of each shot picture in the plurality of shot pictures; determining scene segmentation positions in the plurality of shot pictures according to the fusion semantic information of each shot picture in the plurality of shot pictures, and performing scene segmentation on the plurality of shot pictures. By using the method, the accuracy of video scene segmentation is improved.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

A method for constructing a potential surface anomaly remote sensing detector for single-temporal images

The present invention discloses a method for constructing a potential surface anomaly remote sensing detector for single-phase images. A single-phase remote sensing image is input, and the remote sensing image is subjected to scene segmentation. For each scene in the remote sensing image, its objects are segmented separately, and various spectral characteristics and spatial characteristics of each object are calculated. Then, various spectral characteristics and spatial characteristics of the scene are calculated based on the characteristics of each object in the scene. Then, a difference index between the attributes of each object and the attributes of the scene in which it is located is constructed, namely a potential surface anomaly index. Finally, a suitable threshold is selected to classify the potential surface anomaly index. If the potential surface anomaly index of an object exceeds the threshold, it is determined to be a potential surface anomaly object. Different from existing surface anomaly remote sensing detection methods that require two-phase or multi-phase remote sensing images and can only detect specific surface anomaly events, this method can effectively detect various potential surface anomaly areas in remote sensing images based only on single-phase images. The method is simple in concept and easy to implement. It can be widely used in remote sensing detection of natural disasters, environmental pollution, ecological damage, safety accidents, illegal development and other events, and plays an important early warning role in protecting people's lives and property safety, social security, economic security and political security.
Owner:BEIJING NORMAL UNIVERSITY

Integrated decoder, multi-task network model and training method

This invention discloses an integrated decoder, multi-task network model, and training method for in-vehicle vision systems. The integrated decoder comprises a topological structure composed of a cascade of two-level units, converting the multi-scale feature information output by a base network into a single-scale feature representation. The multi-task network model includes a base network, an integrated decoder, a target detection task head, and a scene segmentation task head. The method employs a balanced convergence method for the multi-task loss function, using empirical values of weight parameters to ensure balanced learning of multiple tasks during backpropagation. This invention improves the accuracy and stability of target detection and scene segmentation algorithms, while also enhancing their real-time performance without compromising their accuracy.
Owner:BEIJING HUAHANG RADIO MEASUREMENT & RES INST

Video scene segmentation method, device, computer equipment, and storage medium

The present application relates to a method, apparatus, computer device, storage medium, and computer program product for segmenting a video scene. The computer device may include a smartphone, a computer, or an intelligent vehicle-mounted device; the method includes: extracting features from a video shot sequence using a dual-path model to obtain shot features; determining positive sample features from the shot features; obtaining negative sample features, and optimizing a first encoder and a second encoder based on the loss value between the negative sample features and the positive sample features; wherein the optimized first encoder serves as a shot feature extraction model; extracting a third shot feature from a target video shot sequence using the shot feature extraction model, and training a scene segmentation model based on the third shot feature; and performing video scene segmentation on the video to be segmented based on the trained scene segmentation model. This method can effectively improve the accuracy of video scene segmentation.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Three-dimensional scene segmentation method and device, equipment and storage medium

The invention provides a three-dimensional scene segmentation method and device, equipment and a storage medium, and relates to the technical field of artificial intelligence and computer vision technologies. The method comprises the following steps: inputting an RGB rendered image into a coder-decoder module, and processing the input RGB rendered image through a two-dimensional visual language model and a coder to obtain Gaussian semantic features; fusing the Gaussian semantic features with other Gaussian attribute features of the three-dimensional Gaussian rendering model to obtain all Gaussian attribute information, inputting the Gaussian attribute information and text information into a large language model, and outputting segmentation feature vectors; inputting the Gaussian semantic feature and the segmentation feature vector into a segmentation information output module to generate a semantic segmentation mask; and based on the semantic segmentation mask, outputting a segmentation result of the target object in the three-dimensional scene. According to the invention, automatic identification and region segmentation of the target object can be realized in a complex three-dimensional environment, so that the intelligent level of a three-dimensional multi-mode interaction and application system is improved.
Owner:SHANGHAI UNIVERSITY OF FINANCE AND ECONOMICS

Road scene segmentation and signal lamp detection equipment

The utility model relates to the field of intelligent traffic, and discloses road scene segmentation and signal lamp detection equipment, which comprises a camera and a base, the bottom of the camera is fixedly connected with a circular column, the circular column is inserted into the base, the outer side of the base is fixedly connected with a shell, the interior of the shell is slidably connected with a push rod, and the push rod is fixedly connected with the base. A hollow ring is fixedly connected to the interior of the shell, a limiting column is fixedly connected to the outer side of the hollow ring, a plug pin is slidably connected to the interior of the shell, a rotating rod is rotatably connected to the outer side of the plug pin, the end, away from the plug pin, of the rotating rod is rotatably connected to the outer side of the push rod, and a limiting rod is fixedly connected to the interior of the shell. According to the utility model, by pushing the push rod, after the plug pin is pulled open, the circular column at the bottom of the camera is inserted into the base, then the push rod is loosened, and the plug pin can be inserted into the circular column under the action of the spring, so that a worker can conveniently disassemble the camera during maintenance or replacement, and the working efficiency is effectively improved.
Owner:马琳琳

A self-supervised video scene boundary detection method based on a timing scene creator

ActiveCN119007085BScene segmentationRadiology
The application discloses a self-supervised video scene boundary detection method based on a timing scene creator, which selects video clips from different pseudo scenes respectively, splices two clips to synthesize a semantic transition point as a pseudo scene boundary. In order to enhance the diversity of the synthesized scene boundary, the application performs shot exchange between the involved video clips. In addition to the pseudo boundary, the application also provides the most likely non-boundary scene through adjacent shots from the same pseudo scene or the shots at the end of the repeated pseudo scene. The application effectively provides high-quality pseudo label data for self-supervised pre-training of video scene segmentation, and significantly improves the accuracy of the video scene segmentation model.
Owner:CHONGQING UNIV

Three-dimensional point cloud segmentation method, system, device and medium

The application provides a three-dimensional point cloud segmentation method, system, device and medium, the method comprising: constructing an initial three-dimensional Gaussian point cloud according to multi-view image data of a target scene; obtaining first base elements located in a boundary blur region in all three-dimensional Gaussian base elements of the initial three-dimensional Gaussian point cloud according to initial semantic masks of each frame of image in the multi-view image data, and performing an adaptive splitting operation on each first base element to obtain a plurality of second base elements; performing multi-scale semantic feature training on a target three-dimensional Gaussian point cloud with the initial semantic mask as a supervision signal to obtain a semantic feature vector of each base element in the target three-dimensional Gaussian point cloud; and assigning a corresponding semantic label to each base element in the target three-dimensional Gaussian point cloud according to the semantic feature vector to obtain a three-dimensional Gaussian point cloud after semantic segmentation. The application realizes optimization of a boundary splitting process of Gaussian splashing and a semantic feature matching process, and improves the accuracy, robustness and flexibility of three-dimensional Gaussian scene segmentation.
Owner:SHANG FEI ZHI NENG JI SHU YOU XIAN GONG SI