Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

40 results about "Scale transformation" patented technology

Multi-scale crowd counting method and system based on VSSM and mask reconstruction, and medium

The invention discloses a multi-scale crowd counting method and system based on VSSM and mask reconstruction and a medium, and belongs to the field of computer vision and deep learning, and the method comprises the steps: carrying out the multi-scale transformation and segmentation operation of an input crowd image, so as to obtain an image block under each scale; analyzing each image block to obtain an information entropy graph, performing masking operation on each image block according to the information entropy graph, inputting the masked image block into a VSSM to perform feature extraction so as to obtain coding features, performing decoding reconstruction on the coding features through a Transform decoder, and finally obtaining a reconstructed feature image; fusing the reconstructed feature images of different scales by using a multi-scale fusion module to generate fused features; and generating a crowd density map based on the fusion features, and analyzing the crowd density map to obtain a crowd counting result. Counting errors caused by single-scale analysis are avoided, and the crowd counting precision in a full-scale range is remarkably improved.
Owner:TIBET UNIVERSITY FOR NATIONALITIES

Lithology generality joint feature classification model construction method

The invention discloses a lithology generality joint feature classification model construction method, which relates to the technical field of intelligent geological prospecting, introduces a cross-scale feature alignment mechanism, and utilizes scale transformation factors to dynamically match the resolutions of macroscopic, mesoscopic and microscopic features, thereby avoiding feature distortion caused by direct fusion; for example, in a volcanic rock and sedimentary rock mixed area, rough texture of a remote sensing image originally conflicts with smooth response of a logging curve, but through alignment processing, the model can automatically balance credibility of different data sources, and misjudgment is reduced; inter-scale dependence modeling is further enhanced through introduction of the feature association graph, vertical sequence association of a thin interbed is captured through a graph structure, the boundary of a millimeter-level thin layer is made clear, and the problem of fuzziness caused by neglecting of interlayer interaction in an existing scheme is solved.
Owner:SHANGHAI LINGYUN INTELLIGENT MINING TECHNOLOGY CO LTD

Gaussian neural field dynamic scene reconstruction system based on depth consistency constraint

The invention provides a Gaussian neural field dynamic scene reconstruction system based on depth consistency constraint, and relates to the technical field of computer graphics, and the system comprises an estimation module which generates a target frame initial depth map; the calculation module reconstructs the point cloud and obtains a point cloud normal direction and a pixel normal direction; the optimization module is used for iteratively correcting the initial depth based on the two types of normal consistency to obtain an optimized depth map; the alignment module is used for determining a scale parameter through regression by taking the first target frame as a reference, and carrying out scale transformation on the depths of other frames to form a consistent depth sequence; the reconstruction module is used for constructing or training a Gaussian neural field based on the sequence and outputting a three-dimensional representation; in addition, the calculation module can contain multi-dimensional wavelets and sparse reconstruction and is used for multi-scale noise suppression and direction weighted fitting. Reference frame selection is based on frame-level quality, geometry and scale stability indexes; according to the system, the intra-frame geometric credibility and the cross-frame scale consistency are improved, ghosting and tearing are reduced, and the stability and integrity of dynamic scene reconstruction are enhanced.
Owner:LISHUI RES INST OF HANGZHOU UNIV OF ELECTRONIC SCI & TECH

Concrete structure damage variable-scale finite element simulation method based on two layers of grids

The invention discloses a concrete structure damage variable-scale finite element simulation method based on two layers of grids, which comprises the following steps: step 1, respectively carrying out structure scale and material scale finite element grid subdivision on a concrete structure to obtain a structure-material two-layer grid covering the whole structure simulation area; 2, endowing material parameters to a structure scale unit and a material scale unit, and respectively establishing a structure scale finite element model and a material scale finite element model of the concrete structure; and 3, based on a simulation scale transformation criterion and a structure scale-material scale connection technology, dynamically updating the structure scale-material scale collaborative finite element model, and simulating a concrete structure damage evolution process. According to the method, the simulation scale of the damage area is converted based on the two layers of grids, so that the simulation precision of the concrete damage evolution process can be ensured, the calculation efficiency is relatively high, and the method has a wide application prospect in the refined analysis of the concrete structure damage evolution process.
Owner:HOHAI UNIV

Lightweight semantic segmentation method and system for remote sensing image of urban scene

The invention provides a lightweight semantic segmentation method and system for a remote sensing image of an urban scene, and relates to the technical field of semantic segmentation, and the method comprises the steps: carrying out the multi-scale transformation of the remote sensing image of the urban scene, carrying out the semantic feature extraction, obtaining a first semantic feature map under each reduced scale, adjusting the first semantic feature map, and obtaining a second semantic feature map under each reduced scale; using a second semantic feature map; performing forward, snake-direction and spiral scanning fusion operation on the second semantic feature map to obtain a third semantic feature map, and fusing the third semantic feature map with the third semantic feature map of the previous scale to obtain a fused semantic feature map; splicing the third semantic feature map and the up-sampling result of the fused semantic feature map of the previous reduced scale according to the channel dimension, and then performing two convolution operations and fusing; obtaining a fused feature image corresponding to a reduced scale; and taking an up-sampling result of the last scaled-down fusion feature image as a semantic segmentation result image.
Owner:TIBET UNIVERSITY FOR NATIONALITIES

Decoupled evolutionary modeling and service recommendation method for user interest state distribution

This invention relates to a decoupled evolutionary modeling and service recommendation method for user interest state distribution. The method first collects user interaction data within a service system to construct a sequence of user-item-time triples. A graph convolution algorithm is used to perform initial static embedding learning of users and items, extracting collaborative signals from the user-item interaction graph. A decoupled state space model is constructed, representing user interest states as distributed variables composed of an orthogonal rotation matrix and a diagonal scaling matrix, respectively capturing the evolution of interest direction and the changing diversity of interest. A user interest precision matrix is ​​generated based on the rotation-scaling transformation, and the anisotropic Mahalanobis distance is used to measure the degree of match between users and items. A recommendation list is constructed based on the distance scores, with the top K best items selected as the recommended results. This method overcomes the static embedding limitations of user interest modeling, achieving invariant preservation and potential preference extrapolation in a dynamically evolving structure, thereby improving the expressiveness and generalization performance of the recommendation system.
Owner:HANGZHOU DIANZI UNIV

A scale transform-based secure numerical multiplication calculation method and device

The application discloses a scale transformation-based secure numerical multiplication calculation method and device, relates to the technical field of secure numerical multiplication calculation, and comprises the following steps: each participant determines a private input value; each participant performs scale transformation on the private input value to generate a private input matrix; each participant takes the private input matrix as input, and calculates a private output matrix by using a secure two-party matrix multiplication protocol based on a secure data confusion technology; each participant calculates the sum of each element of the private output matrix to obtain a private output value; and each participant sends the private output value to a calculation requester to obtain a secure numerical multiplication calculation result. By introducing scale transformation and the secure two-party matrix multiplication protocol based on the secure data confusion technology, the application can improve the calculation efficiency, the communication efficiency and the calculation accuracy, and does not depend on a third-party cloud platform, thereby further guaranteeing security.
Owner:BEIHANG UNIV

A target detection method, device and medium based on multi-scale morphological operators

ActiveCN120032135BCharacter and pattern recognitionImaging processingStructuring element
The present invention provides an object detection method, device and medium based on multi-scale morphological operators, belonging to the field of object detection in image processing. It includes: processing a single-frame image to obtain a preprocessed image; performing multi-scale morphological dilation operation pixel by pixel to obtain the dilation result and erosion result of each pixel under each scale structuring element; constructing a gray-scale transformation index using the dilation result and erosion result, and marking the scale of the structuring element with the smallest gray-scale transformation index value as the optimal scale; calculating the mean value of the pixel points corresponding to the optimal scale structuring element; for each pixel point, constructing a morphological feature value using the mean value, dilation result, and the gray value of the pixel point, and determining whether the pixel point is a target. The multi-scale morphological operation of the present invention has the characteristics of high detection probability and low false alarm compared with traditional methods, and can effectively alleviate false alarm events caused by local gray-scale changes.
Owner:INST OF OPTICS & ELECTRONICS CHINESE ACAD OF SCI

Decoupling evolution modeling and service recommendation method oriented to user interest state distribution

The invention relates to a user interest state distribution-oriented decoupling evolution modeling and service recommendation method, which comprises the following steps of: firstly, collecting interaction data of a user in a service system, and constructing a user-article-time triple sequence; performing initial static embedding learning on the user and the article based on a graph convolution algorithm, and extracting a cooperative signal in the user-article interaction graph; constructing a decoupling state space model, representing a user interest state as a distribution variable composed of an orthogonal rotation matrix and a diagonal scaling matrix, and respectively capturing interest direction evolution and diversity change characteristics; generating a user interest precision matrix according to rotation-stretching transformation, and measuring the matching degree between a user and an article by adopting an anisotropic mahalanobis distance; and constructing a recommendation list according to the distance scores, and taking the first K optimal articles as recommendation results. According to the method, static embedding limitation of user interest modeling is broken through, invariant keeping and potential preference extrapolation in a dynamic evolution structure are achieved, and the expression ability and generalization performance of a recommendation system are improved.
Owner:HANGZHOU DIANZI UNIV

Information reasoning method, device and computer equipment

The present invention relates to the field of natural language processing, and in particular to an information reasoning method, apparatus, computer equipment, and storage medium, which perform multi-scale transformation on a knowledge graph, construct a multi-scale zoom graph of the knowledge graph, and perform multi-scale interaction in combination with a subgraph to be inferred and a multi-scale transformation weight matrix constructed based on the subgraph to be inferred, aggregate neighboring node information from different scales, make full use of information at different scales, and construct a more accurate retrieval subgraph for performing graph information reasoning, thereby improving the accuracy of accurate information reasoning on the graph.
Owner:SHANGHAI YUANYUAN INFORMATION TECH CO LTD

Secure real-number multiplication method and apparatus based on scale transformation

Provided is a secure real-number multiplication method and apparatus based on scale transformation, which relates to the technical field of secure real-number multiplication. The secure real-number multiplication method includes: determining, by each participant, a private input value; performing, by each participant, scale transformation on the private input value to generate a private input matrix; with the private input matrix as input, computing, by each participant, a private output matrix by utilizing a Secure Two-Party Matrix Multiplication (S2PM) protocol based on a secure data disguising technology; calculating, by each participant, a sum of all elements in the private output matrix to obtain a private output value; and sending, by each participant, the private output value to a computation requester to obtain a secure real-number multiplication result.
Owner:BEIHANG UNIV

A tongue feature recognition method and device, electronic equipment and storage medium

Embodiments of the present application disclose a tongue feature recognition method and device, electronic equipment and a storage medium. The method comprises: acquiring a tongue image; performing multi-scale operation on the tongue image to determine at least two scale transformation images; inputting the at least two scale transformation images into a recognition model to output target feature information corresponding to the tongue image through the recognition model; and the recognition model is a cross-network model composed of a segmented convolutional neural network (Seg-CNN) structure and a point convolutional neural network (Point-CNN) structure. In this way, when analyzing the tongue image, the multi-scale operation is performed on the tongue image, and the network model composed of the Seg-CNN structure and the Point-CNN structure is used, so that the global feature information with points, lines and surfaces can be recognized, thereby the crack feature in the tongue image can be clearly recognized, and the recognition accuracy of the crack feature is improved.
Owner:CHINA MOBILE CHENGDU INFORMATION & TELECOMM TECH CO LTD +1

Multi-scale test evaluation method for pavement performance of steel bridge deck

The invention provides a steel bridge deck pavement performance multi-scale test evaluation method which comprises the following steps: acquiring test data of steel bridge deck pavement, and extracting a multi-scale performance parameter set from the test data; based on the multi-scale performance parameter set, establishing a scale conversion function for mapping local performance parameters to an overall structure state; constructing an initial digital model based on the scale transfer function and the multi-scale performance parameter set, and performing parameter reverse calibration on the initial digital model in combination with actually measured response data in the full-scale segment scale data to obtain a multi-scale performance digital model; and based on the multi-scale performance digital model, simulating the obtained service environment data of the target pavement structure to obtain performance evaluation data of the target pavement structure, the performance evaluation data including a predicted performance degradation value and a predicted remaining service life. By the adoption of the method, more reliable technical support can be provided for performance management and control of steel bridge deck pavement.
Owner:HUBEI ROAD & BRIDGE GRP CO LTD

Real-time video jitter removal method based on scale change

The invention relates to the technical field of image processing, and discloses a scale change-based real-time video jitter removal method, which comprises the following steps of: 1, acquiring a video frame at a previous moment as a pre-frame through single video shooting equipment, and acquiring a video frame at a current moment as a processing frame; step 2, pre-processing the obtained pre-frame and the processed frame; and step 3, extracting a preset number of feature points in the pre-processed pre-frame and the processed frame, matching the obtained feature points, removing mismatching points by using an RANSAC algorithm, and performing scale transformation and mapping on the matched feature points to the original size to obtain matching points of the pre-frame and the processed frame under the original size. The method comprises the following steps: preprocessing an acquired pre-frame and a processed frame, generating a low-resolution image through scale transformation downsampling, extracting feature points from the low-resolution image, removing mismatching points, and mapping the feature points to matching points with original sizes.
Owner:ANHUI CIVIO INFORMATION & TECH

Secure two-party data comparison method and apparatus based on scale transformation

Provided are a secure two-party data comparison method and apparatus based on scale transformation. The method includes: transmitting, by a computation requesting party, a two-party data comparison request to two participant nodes; performing, by each participant node, scale transformation and linear scaling on private data locally after receiving the two-party data comparison request, to obtain an encrypted vector; determining, by each participant node, a real number locally using a secure two-party dot product protocol based on the local encrypted vector, and sharing the real number with the other participant node; determining, by each participant node, a comparison sign based on the obtained real number and transmitting the comparison sign to the computation requesting party; and determining, by the computation requesting party, a comparison result of the private data of the two participant nodes based on the comparison signs transmitted by the two participant nodes.
Owner:BEIHANG UNIV

A point cloud shape analysis method based on spatial geometry perception convolutional neural network

The application relates to the field of 3D vision, in particular to a point cloud shape analysis method based on a space geometry perception convolutional neural network. According to a K nearest neighbor algorithm, a point cloud is constructed as graph structure data, and a domain adaptive convolution kernel with different geometric shapes is adaptively generated according to features in each neighborhood. Subsequently, unit direction vectors, elevation angles and azimuth angles of neighborhood nodes are taken as prior geometric information, and convolution operation is performed on the generated convolution kernel to realize space geometry perception convolution operation. Subsequently, graph attention pooling is designed to coarsen the point cloud, realize multi-scale analysis and reduce the calculation cost. Finally, space geometry perception convolution and graph attention pooling are taken as basic units to construct two networks to realize point cloud classification and component segmentation tasks. The application can effectively classify and component segment target point clouds, and ensures the invariance of point cloud translation and scaling transformation, so that the model has stronger robustness.
Owner:SHANXI UNIV

RGB-D salient target detection method based on texture enhancement guidance

The invention discloses an RGB-D salient target detection method based on texture enhancement guidance. The RGB-D salient target detection method is specifically implemented according to the following steps: step 1, constructing a data set and an encoder; step 2, constructing a texture enhancement module; step 3, constructing a dual-path adaptive interaction module; and 4, constructing a dynamic decoding module. In order to solve the key problems of heterogeneous modal feature degradation, depth noise cross-layer propagation, multi-scale semantic mismatch and the like in the prior art, noise suppression of depth features is realized by using high-frequency texture prior constraints, and cross-modal semantic association is established through a channel-space collaborative dynamic interaction mechanism. A deformable cross-scale transformation technology is adopted to realize progressive calibration of multi-level features, and the method shows remarkable boundary integrity and noise suppression advantages in input scenes of multiple targets, low-quality depth and the like.
Owner:XIAN UNIV OF TECH

Tax-related legal text-oriented named entity recognition dependent enhancement method

The application discloses a kind of tax-related legal text-oriented named entity recognition dependent enhancement method, comprising: tax named entity recognition is regarded as span classification task, a large number of spans are enumerated from input text by sliding window, and the deep representation of each span is generated by feature splicing method;Introduce a contrastive learning loss, and the contrast relationship is mined from highly overlapping span;Scale transformation mechanism is used to realize span interaction, and the geometric information of each candidate span is embedded in native span representation, to encode the interaction dependency between spans.The application converts tax named entity recognition into span classification task, and fully mines the interaction dependency between entities, realizes strong inference relationship, and introduces contrastive learning to improve the discrimination degree between different types of highly overlapping entities, so that the named entity in tax legal text can be more accurately and reasonably recognized, laying a foundation for downstream tasks such as tax incentives.
Owner:XI AN JIAOTONG UNIV

Information extraction method and device, computer equipment and storage medium

The invention relates to the field of information extraction, and provides an information extraction method, device and equipment and a computer storage medium, and the method comprises the steps: obtaining a to-be-processed image, and carrying out the scale transformation of the to-be-processed image, and obtaining a to-be-enhanced image of at least one scale; inputting each to-be-enhanced image into a corresponding feature extraction network to obtain a feature image corresponding to the to-be-enhanced image of each scale; performing feature fusion on each feature image to obtain an enhanced image of an original scale; performing reconstruction processing on the enhanced image based on the to-be-processed image to obtain a target image; and based on a preset information extraction vector, obtaining target information matched with the information extraction vector in the target image.
Owner:CHINA PING AN LIFE INSURANCE CO LTD

Three-dimensional reconstruction method and system based on multi-view image fusion

The invention discloses a three-dimensional reconstruction method and system based on multi-view image fusion. The method comprises the following steps: carrying out adaptive segmentation on an obtained long image line by line to generate a segmentation mask; executing cylindrical surface expansion mapping on the target surface according to the mask segmentation result to obtain corrected plane textures; seamless fusion processing is carried out on the plane textures under the multiple visual angles to generate 360-degree panoramic textures; based on the pixel width W of the panoramic texture, performing scale transformation and displacement processing on the coordinates of the texture U; and mapping the panoramic texture subjected to scale transformation and displacement processing on the texture U coordinates to a three-dimensional cylindrical model to generate a high-fidelity three-dimensional reconstruction model of the target object. According to the method, the problems of insufficient image segmentation precision, incomplete cylinder perspective distortion correction, visual flaws of multi-view texture splicing and joints in three-dimensional model rendering in the three-dimensional reconstruction process of a slender cylindrical object (such as a steel wire rope) can be effectively solved.
Owner:XUZHOU SUNWELL MINING TECH CO LTD +1

A distance scale transformation method

The application provides a range scale conversion algorithm, and belongs to the technical field of airborne SAR polar format algorithm (PFA) real-time imaging, and specifically comprises the following steps: embedding compensation processing of a fast time offset Delta tau corresponding to each pulse into a matched filter to complete scene center point motion compensation and pulse compression processing of each pulse; and then performing range intercept processing on the pulse compressed data to complete imaging processing of a region of interest near an imaging center, so that the problem of image winding caused by range intercept in the original algorithm can be solved, and the azimuth interpolation processing time and the storage capacity requirement of an imaging system are greatly reduced.
Owner:LEIHUA ELECTRONICS TECH RES INST AVIATION IND OF CHINA

Table line detection method, device, computer device and storage medium

The present application relates to a table line detection method, apparatus, computer equipment, storage medium, and computer program product, which can be used in the financial field or other fields. The method comprises: acquiring a table line image, performing a multi-scale transformation on the table line image to obtain a multi-scale image set, obtaining a fused table line image based on the multi-scale image set through region growing and region fusion, and extracting the table lines from the fused table line image through a straight line detection algorithm. By adopting this method, the multi-scale image set obtained by performing a multi-scale transformation on the table line image can show clearer table line features. By performing region growing and region fusion, the multi-scale image set with different features can be fused. This method of extracting table lines based on the straight line detection algorithm combined with multi-scale transformation can improve the accuracy of table line detection.
Owner:INDUSTRIAL AND COMMERCIAL BANK OF CHINA

Method, system, medium and program product for measuring critical exponents of quantum phase transitions

The present disclosure relates to a method, a system, a medium, and a program product for measuring critical exponents of quantum phase transitions, and relates to the field of quantum computing and quantum simulation. The method includes: for each candidate value of the critical exponent, determining multiple value combinations of the change rate of the system parameters and the system size of the quantum system such that the candidate value, the change rate, and the system size satisfy a first condition; for each value combination of the change rate and the system size: linearly changing the system parameters at the change rate such that the quantum system undergoes a quantum phase transition; measuring a first correlation function of the quantum system; performing a scale transformation on the first correlation function that satisfies a second condition to obtain a second correlation function; for each candidate value, comparing the differences between the second correlation functions; and selecting the candidate value that minimizes the difference as the measurement result of the critical exponent. This method can overcome the problems of short coherence time and limited system size of the quantum system and improve the measurement accuracy.
Owner:LIANGYI WANXIANG (BEIJING) TECHNOLOGY CO LTD

Seismic event detection method based on multi-scale transformation of seabed application scene

The invention discloses a multi-scale transformation seismic event detection method based on a seabed application scene, and relates to the technical field of seismic event detection. According to the seismic event detection method based on the multi-scale comprehensive transformation of the seabed application scene, a plurality of channel original signals are collected through a seabed seismograph, and a fusion feature vector is constructed through multi-scale time-frequency transformation; performing local event preliminary screening, generating an event packet, and reporting the event packet to the buoy node; in the buoy nodes, event packets uploaded by all the ocean bottom seismometers in the jurisdiction range of the buoy nodes are received, and a cluster report is generated and transmitted to the shore-based processing center after the weighted comprehensive confidence coefficient is analyzed; in a shore-based processing center, cluster event reports uploaded by all buoy nodes within the range where the shore-based processing center is located are received, and the global existence probability of the same event is analyzed. And determining the same event higher than a preset global existence probability threshold as a real seismic event.
Owner:CHONGQING GEOLOGICAL INSTR FACTORY

Method and system for detecting large language model LoRA fine tuning origin

The invention relates to the technical field of large language models, in particular to a method for detecting a large language model LoRA fine tuning origin. Comprising the steps of generating an adaptability vocabulary, recording intermediate features of a to-be-verified model, selecting a base candidate model, obtaining output features of the base candidate model, calculating approximate intermediate features of the base model, extracting LoRA rank information through singular value decomposition, determining minimum rank information and judging a fine tuning origin. According to the method and the system for detecting the LoRA fine-tuning origin of the large language model, the fine-tuning origin of the model can still be accurately detected in the face of confusion technologies such as parameter replacement and zoom transformation, the defect of confusion resistance in the prior art is effectively overcome, the LoRA rank information used in the fine-tuning process can be accurately extracted, a detailed basis is provided for model verification, and the method and the system are suitable for popularization and application. The method facilitates further analysis of fine adjustment details of the model, is suitable for large language models of various architectures and scales, is not limited by the size of the model and specific fine adjustment parameters, and has wide applicability.
Owner:SHANGHAI JIAOTONG UNIV

Three-dimensional character model appearance transformation method, device, equipment, medium and product

The present invention discloses a method, device, equipment, medium, and product for transforming the appearance of a three-dimensional character model, relating to the technical field of character appearance transformation. The method comprises the following steps: obtaining a target three-dimensional character; mapping the skeleton of the target three-dimensional character onto a standard human skeleton to obtain a first skeleton of the target character; determining a second skeleton of the target character based on the first skeleton; the second skeleton of the target character being a skeleton obtained by adjusting the first skeleton of the target character to a T-Pose posture; performing bone rotation and scaling transformations on the second skeleton of the target character to obtain an initial binding posture; determining a deformation template; determining a deformation part of the target three-dimensional character based on the initial binding posture and the deformation template; and determining adjustment parameters for the skin and skeleton of the deformation part based on the deformation parameters of the deformation template. The present invention can easily change the appearance of a character by applying the deformation template.
Owner:JILIN ANIMATION INST +1

Building polygon simplification method in map synthesis based on large language model normal form

The invention discloses a method for simplifying a building polygon in map synthesis based on a large language model normal form. The method comprises the following steps: acquiring a coordinate data sequence of a vector building polygon to be simplified; the coordinate data of the vector building polygon to be simplified are input into a simplification model based on a large language model normal form, a simplified polygon is obtained, simplification of the building polygon in scale transformation from a large scale to a small scale is achieved, and the simplification model constructs a vocabulary through a gridding mechanism; discretizing an original coordinate sequence in the vector building polygon to be simplified into a Token sequence; a mask self-attention mechanism of a Transform of a decoder structure is adopted to learn context dependence of a polygon point sequence, and feature embedding of each point is formed; and finally, point-by-point generation of the simplified polygon is realized through a full-connection neural network.
Owner:HUAZHONG NORMAL UNIV

Action recognition methods, systems, devices, and media based on spatiotemporal scale transformation

This invention discloses a behavior recognition method based on spatiotemporal scale transformation, comprising: detecting a target using a target detector to obtain corresponding target bounding box regions; performing spatial multi-scale transformation on each target region to obtain multi-scale target regions; fusing features of the multi-scale target regions to obtain fused features; and classifying and recognizing behaviors based on the fused features. This invention combines target detection with behavior recognition, enabling the behavior recognition model to effectively focus on the regions where behaviors actually occur. Addressing the problem of inconsistent target spatial scales, this invention utilizes upsampling and downsampling to achieve scale uniformity, better ensuring the scale consistency of targets. To address the temporal differences between different samples, this method proposes a multi-frame-rate sampling operation to fuse features from multiple time scales, enabling effective feature extraction for behaviors of different durations.
Owner:SHANGHAI JIAOTONG UNIV

Infrared and visible light image fusion method fusing adaptive decomposition and multi-scale transformation

The invention discloses an infrared and visible light image fusion method fusing adaptive decomposition and multi-scale transformation, and relates to the technical field of multi-modal image fusion. In order to solve the problems of detail loss, noise interference, structural distortion and the like in infrared and visible light image fusion, the method comprises the steps of firstly performing normalization and registration preprocessing on an input image, then extracting low-frequency and high-frequency features by adopting adaptive Fourier decomposition, and performing sparse representation and orthogonal matching pursuit fusion on a low-frequency part to obtain a low-frequency part and a high-frequency part; and the high-frequency part guides dynamic weighted fusion according to local energy and information entropy, and finally reconstruction is performed and cross-scale consistency verification is introduced to improve fusion robustness. The method can be widely applied to fusion of infrared and visible light images, and the quality of the fused image is improved.
Owner:EAST CHINA UNIV OF SCI & TECH

A highly fault-tolerant multimodal data fusion method based on paired feature synthesis network

The present invention discloses a highly fault-tolerant multimodal data fusion method based on a paired feature synthesis network. The method specifically comprises: dynamically assigning weights to the input raw data and preprocessing it to obtain initial refined feature pairs; performing scale transformation on the initial refined feature pairs to obtain input features; and classifying and predicting the input features to obtain multimodal data fusion classification results. By dynamically assigning weights to the input raw data and preprocessing it, the present invention ensures the system's fault tolerance and reliability when a single data input fails. Furthermore, through the scale transformation operation, spatial information can be highlighted, the feature representation of edge regions can be enhanced, the classification accuracy of the algorithm can be improved, and the waste of model parameters can be reduced. The method can be widely applied to the field of computer vision multimodal data fusion technology.
Owner:SUN YAT SEN UNIV