Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

84 results about "Visual modality" patented technology

Visual Modality - A Visual Learner. Learns by seeing and by watching demonstrations Likes visual stimuli such as pictures, slides, graphs, demonstrations, etc. Conjures up the image of a form by seeing it in the “mind’s eye” Often has a vivid imagination.

Multi-modal dialogue emotion recognition method and system based on cross-modal fusion and comparative learning

The invention relates to the technical field of natural language processing, in particular to a multi-modal dialogue emotion recognition method and system based on cross-modal fusion and comparative learning. According to the method, learnable residual scaling and pre-normalization are introduced through a cross-modal encoder, deep interaction of texts, voices and visual modals is stabilized, and gradient explosion is inhibited; a dialogue graph fusing semantic similarity and time proximity is constructed online through a semantic-time sequence graph enhancement module, and a graph attention network is used for explicitly modeling long-distance round dependence and cross-speaker emotion transmission; through an adaptive comparison and alignment module, a dynamic scheduling comparison loss and index moving average updated mode-emotion prototype library is adopted to realize data distribution adaptive cross-mode alignment; through cooperative work of the modules, the problems that in the prior art, cross-modal fusion is unstable, long-distance and cross-speaker dependence modeling is insufficient, and cross-dataset alignment capacity is weak are solved.
Owner:CHONGQING TELECOMM PLAN & DESIGN INST

Textile cloth defect real-time detection method and system based on multi-modal feature fusion

The invention provides a textile cloth defect real-time detection method and system based on multi-modal feature fusion, and the method comprises the steps: collecting visual image data, infrared thermal imaging data and ultrasonic acoustic data of textile cloth through a multi-sensor array, and forming multi-modal input; performing time sequence alignment and noise filtering preprocessing on the multi-modal data to eliminate motion artifacts and environmental interference; a parallel feature extraction module is used for extracting texture features from the visual data, extracting temperature distribution features from the thermal imaging data and extracting acoustic impedance features from the acoustic data. By deeply fusing complementary information of three modes of vision, thermal imaging and ultrasonic wave, the detection capability is improved, the vision mode captures surface texture details, the thermal imaging mode reveals thermodynamic anomalies related to friction and materials, the ultrasonic wave mode perceives subcutaneous structure defects, hidden flaws which cannot be recognized by a single mode can be found, and the detection efficiency is improved. Therefore, the omission ratio is greatly reduced, and flaw types are distinguished more accurately.
Owner:NANCHANG ZHONGTUO KNITWEAR CORP LTD

Video segmentation method, server, storage medium, and program product

The present application provides a video segmentation method, a server, a storage medium, and a program product. In the method of the present application, video data to be segmented is segmented into multiple data segments, unimodal features of the data segments, including text features of a text modality and visual features of a visual modality, are respectively extracted by means of a video topics segmentation model, and then the text features and visual features of the data segments are fused, so that the fusion of multimodal information can be performed at the intermediate representation level, the relationship and interaction between different modalities can be better captured, and higher-quality multimodal fusion features of the data segments are obtained. Furthermore, on the basis of the multimodal fusion features of the data segments, whether the data segments are topic boundaries is predicted, so that the topic boundaries of the video data can be accurately predicted, improving the accuracy of topic boundary recognition, thereby improving the accuracy and quality of video topics segmentation results.
Owner:ALIBABA (CHINA) CO LTD

Multi-modal sentiment analysis model based on multi-granularity features and adaptive fusion

The invention discloses a multi-modal sentiment analysis model based on hierarchical adaptive cross-modal fusion, belongs to the field of natural language processing, and is used for solving the problems that in the prior art, multi-granularity sentiment feature extraction is insufficient, a cross-modal fusion mechanism is rigid, and the distribution difference between different-source modals is large. The method comprises the following steps: firstly, extracting features of texts, audios and visual modalities from original video data, and coding the features into advanced semantic features; secondly, multi-granularity information is fused through a hierarchical feature extractor to generate enhanced single-mode features; then, a self-adaptive cross-modal fusion network with a text as a core is adopted to realize bidirectional interaction and dynamic weighted fusion between modals; further, a dynamic contrast learning mechanism is introduced to align modal distribution in a unified potential space; and finally, inputting the optimized multi-modal features into a classifier and outputting an emotion analysis result. According to the model, through collaborative optimization of multi-granularity feature extraction, adaptive fusion and comparative learning, the accuracy of sentiment analysis and the robustness of the model are remarkably improved.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Label-assisted report generation method and device

The invention relates to a method and a device for generating a report under the assistance of a label, and the method comprises the following steps: 1) extracting a structured label set from a text report of a sample based on a large language model, wherein the structured label set comprises a multi-classification group consisting of dichotomous labels and mutual exclusion options; 2) aggregating the labels in batches, after a threshold value is reached, merging and de-weighting, performing specification and mutual exclusion group merging on synonymous, near-synonymous and redundant labels, and converging into a unified label library; 3) based on the text report and the tag library, outputting a tag subset of each sample through a large language model; 4) multi-modal multi-label classification model training: extracting each visual modal feature, performing weighted aggregation and splicing, and outputting each label group logits through a classification head to perform weighted group loss optimization; (5) carrying out joint training on the multi-modal large language model by using samples of'only images-reports' and'images + labels-reports', and (6) carrying out label prediction and screening on the images by using the classification model, and inputting'images + prediction labels' into the multi-modal large language model to obtain a final report.
Owner:ZHEJIANG UNIV

Adversarial camouflage and decoy generation across visual and non-visual modalities

Disclosed is a system to generate an adversarial pattern. The system receives an input specifying data describing a target object, an objective indicating whether to reduce detectability of the target object or to induce a false detection of a target type, and context parameters. The system generates candidate adversarial patterns for the target object using an AI-based generative algorithm. Each candidate adversarial pattern represents a potential solution for achieving the objective under context parameters. The system simulates a representation of the target object with candidate adversarial pattern in a virtual environment that models a plurality of sensor modalities. The simulation may use machine learning object detection models. The system analyzes the target object and generates, for each target object, a performance metric for ranking the candidate adversarial pattern based on an effectiveness of the simulated representation. The candidate adversarial pattern with highest rank is generated as an optimized adversarial pattern.
Owner:THE ADVERSARIAL CO INC

System

A system is provided.SOLUTION: A system comprising: means for storing information; means for extracting key claims from the information using generative artificial intelligence to analyze the information; means for quantifying the key claims; and means for visually displaying the quantified key claims.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

A multi-modal entity linking method based on double encoders and hybrid expert mechanism

A multimodal entity linking method based on dual encoders and a hybrid expert mechanism is proposed. This invention relates to multimodal entity linking technology at the intersection of natural language processing and computer vision. Addressing the problems of low inference efficiency, insufficient cross-modal interaction, and shallow modal fusion in existing methods, this invention proposes a multimodal entity linking method based on dual encoders and a hybrid expert mechanism. A dual-tower architecture is used to independently encode mentions and entities. Entity embeddings can be pre-computed offline and indexed, and linking is completed during inference through fast vector retrieval. A hybrid expert mechanism is introduced to achieve adaptive feature transformation of samples, and a gating network dynamically selects expert combinations. Bidirectional cross-modal attention is used to establish fine-grained alignment at the word-image block granularity. A channel attention mechanism dynamically balances the contributions of textual and visual modalities. The model is jointly optimized by multiple constraints, including load balancing loss. This invention achieves efficient retrieval while maintaining deep inference capabilities, simplifies inference time complexity, and is suitable for scenarios such as knowledge graph construction and intelligent question answering systems.
Owner:CHINA UNIV OF PETROLEUM (EAST CHINA)

A parkinson's disease early identification system and method based on multi-task learning

The application provides a Parkinson disease early identification system and method based on multi-task learning, and belongs to the technical field of machine learning. First, a Parkinson disease discrimination data set covering action, language and visual modalities is constructed to provide comprehensive symptom data and identity for subsequent analysis. Then, in combination with medical prior rules, a symptom recognition module integrating multiple traditional machine learning models is constructed to preliminarily discriminate the action and language data. A short-term recognition model is constructed to enhance the multi-modal disease perception ability and improve the discrimination effect of short-term Parkinson disease recognition. Finally, based on the recognition result sequence of multiple time periods, mathematical modeling analysis is carried out to fuse statistical trends and group differences, and to realize stable comprehensive judgment of Parkinson disease. Through the collaborative training and knowledge sharing among tasks, the model can enhance the comprehensive discrimination ability of multi-modal pathological signals of Parkinson disease, so as to realize more accurate and sensitive early identification.
Owner:YANTAI LIAN BIOTECHNOLOGY CO LTD

system

We provide the system. [Solution] A means of analyzing electronic communications and determining priorities, Means for generating notifications based on the aforementioned priority, A means of automatically generating replies according to criteria specified by the user, A means for sending the automatically generated reply, An information display device provides means for providing notifications by sound and visual means, A system that includes this.
Owner:SOFTBANK GROUP CORP

Video generation method for generating video showing air permeability of constituent element of wearable article

Provided is a video generation method for generating a video showing the air permeability of a constituent element of a wearable article. The excellent air permeability of the components of the wearable article is visually shown so that the wearable article can be easily recognized by a user. The method comprises the following steps: a mounting step, a part of a constituent element of a wearable article (for example, a waist strip element having improved air permeability when stretched) is attached to a water vapor discharge device for discharging predetermined visually recognizable water vapor in a stretched state (for example, in a range from 60% of the maximum elongation to the maximum elongation) in which the actual use state is reproduced; and an imaging step for imaging, in a specific discharge direction (for example, + / -20 degrees in the vertical direction or 45 degrees in the horizontal-oblique upper direction), a case in which the water vapor discharged from the water vapor discharge device passes through the constituent element (for example, a state in which the water vapor rises linearly or rises upward by 20 cm or more).
Owner:UNI CHARM CORP

College student behavior anomaly early warning method and system based on multi-source data fusion

PendingCN122332841AImage manipulationNoise
This invention discloses a method and system for early warning of abnormal behavior among college students based on multi-source data fusion, relating to the technical field of image processing. The method involves real-time acquisition of video streams of individuals within a target area, extraction of behavioral features from the video streams for classification, and a determination list based on these behavioral features. Multi-source data is collected, and anomaly coefficients are calculated based on the multi-source data. If the anomaly coefficient exceeds a threshold, the target individual is identified as exhibiting abnormal behavior. Anomaly warnings are then issued based on the behavioral features corresponding to the abnormal behavior. By calculating the anomaly coefficient through multi-source data fusion, a preliminary screening of hidden risks is achieved, reducing the detection delay and computational redundancy caused by traditional full-process monitoring analysis. Furthermore, real-time video stream acquisition and cloud-based behavioral feature extraction are initiated for high-risk individuals, avoiding the limitations of audio noise interference and single visual modalities, and improving the robustness of identification and the real-time performance of warnings in complex monitoring scenarios.
Owner:ZHEJIANG GUGEL INTELLIGENT TECHNOLOGY CO LTD

Unsupervised learning defect detection method based on few sample data

The invention belongs to the technical field of few-sample data, and discloses an unsupervised learning defect detection method based on few-sample data, comprising the following steps: step 1, acquiring a few-sample normal sample and corresponding physical attribute data, the normal sample being a defect-free target object image, and constructing a heterogeneous feature extraction network to obtain a feature extraction result; a visual mode and a physical attribute mode are deeply fused, a visual branch extracts visual features such as textures and edges through MobileNet, a physical branch captures physical features such as surface roughness and temperature distribution through Transformer, and interactive fusion and unified mapping of the two features are achieved by means of a cross attention mechanism. By means of the design, multi-dimensional characteristics of hidden defects such as metal hidden cracks and uneven density in glass can be accurately captured, the defects have no obvious difference visually, but can cause physical attribute abnormity, and characteristic information of the defects can be completely reserved through heterogeneous characteristic vectors.
Owner:CHANGZHOU MICROINTELLIGENCE CO LTD

system

We provide the system. [Solution] A means of accessing an educational database and obtaining data for the subject being studied, A means of analyzing acquired data and automatically generating quiz questions, A means for selecting images or diagrams related to the generated quiz questions and integrating them visually, A means of sending integrated quizzes and visuals to the user's terminal, A means of collecting and analyzing user responses, A means of suggesting the next learning content based on the user's learning progress, A system that includes this.
Owner:SOFTBANK GROUP CORP

Behavioral Change Visualization Card

ActiveJP1829704SAlgorithmBrain section
The cards related to this design are diagram cards that visually represent the behavioral change process corresponding to neural activity by combining color bands based on five colors: black, green, red, yellow, and blue, with geometric and character elements that symbolize growth. The overall composition is unified in the order of black → green → red → yellow → blue, and the upper left logo, upper right character, and brain diagram are also composed based on the same color scheme rule. [Upper left logo structure] A logo mark resembling a plant is placed in the upper left. This logo is a trademarked design and is an element that symbolizes the image of growth and the natural cycle. The red → yellow → blue color scheme within the logo represents a transition of colors that expresses gradual change and visually corresponds to the color band of the entire card (black → green → red → yellow → blue). [Upper right structure] Multiple characters are placed in the upper right, and the stages of the behavioral change process corresponding to neural activity are expressed by the red, yellow, and blue color scheme that is consistent with the overall color band. The brain diagram in the upper right is also composed of four colors: green, red, yellow, and blue, and in accordance with the color scheme rules of the entire card, it is a diagram element that visually represents brain regions corresponding to neural activity. [Central Structure] The central part of the card has a structure based on the shape of the number "8," and the lower (1-7), middle (9-15), and upper sections are arranged in an infinity symbol (∞) shape. This ∞ structure is an element that visually represents the cyclical relationship between the foundational formation stage (1-7) and the behavioral change stage (9-15) corresponding to neural activity, and is a characteristic composition of this design. [Color and Unity] This design is unique in that the following elements are unified based on a color band of black → green → red → yellow → blue. ● Upper left logo ● Upper right character and brain diagram ● Growth lines arranged in the upper, middle, and lower sections ● Overall five-color band structure These unified color schemes make it possible to visually represent the processes, growth stages, and behavioral stages corresponding to neural activity in a continuous manner. [Flow indicated by arrows] Multiple arrows connecting each stage indicate a direction that aligns with the flow of the five color bands, making it easy for users to visually follow the transitions between stages. The placement of the arrows is an important visual element that reinforces the continuity and unity of the entire card.
Owner:手島 正太

System

An object of the system according to the embodiment is to convey a task to a child who is not good at understanding words in a visually easy-to-understand manner.SOLUTION: A system includes an task management unit, a visualization unit, and a completion notifying unit. The task management unit manages a task list. The visualization unit visually displays the tasks managed by the task management unit. The visualization unit generates an image according to the content of the task using the generation AI. The completion notification unit visually expresses the result of the task when the child completes the task.SELECTED DRAWING: Figure 1
Owner:SOFTBANK GROUP CORP

Multi-modal generation type pre-training method and system and image question and answer generation method

The invention relates to a multi-modal generation type pre-training method and system and an image question and answer generation method. Visual modal data, language modal data and multi-modal data are obtained; training a visual base model on the visual modal data through a self-supervised comparative learning normal form, and storing a prediction layer weight as a visual prototype; a language base model is trained on the language modal data through a self-supervised lexical element prediction normal form, and word embedding weights are saved to serve as language prototypes; bridging the trained visual base model and language base model, and obtaining lexical element prediction loss through a multi-modal lexical element prediction normal form; cross-modal knowledge distillation and self-modal knowledge distillation are carried out on multi-modal data, and supervised pre-training is carried out. Compared with the prior art, the method has the advantages that the prior information of the visual modality and the language modality can be effectively utilized, and high-precision understanding and reasoning in a cross-modal scene are realized through unified multi-modal generation type pre-training.
Owner:FUDAN UNIVERSITY

Context compression method and system based on visual modality and medium

The invention provides a context compression method and system based on a visual modality and a medium, and the method comprises the steps: obtaining long text context data, and rendering the long text context data into a corresponding document image; performing context optical coding on the document image, and performing segmentation processing on the document image by adopting a visual segmentation model to obtain a plurality of initial visual tokens; performing down-sampling compression on the initial visual token based on a convolution compression algorithm to obtain a compressed token; capturing a long-distance dependency relationship among different compression tokens based on a global attention mechanism, extracting high-level semantic knowledge, and outputting a final visual token sequence; executing context optical decoding on the final visual token sequence, and fusing context optical decoding data with a text prompt input by a user to generate a target text; through visual modal compression, computing resources and memory occupation required for processing a long text are reduced, and relatively high text decoding precision can still be maintained while a high compression ratio is maintained.
Owner:SHENZHEN KUSI BIOTECHNOLOGY CO LTD

Storage device

To provide a storage device that allows visual confirmation that a stored item exceeds its load capacity. [Solution] The storage device 10 comprises a base part 20 having a fixed part 21 fixed to the left wall surface W1, a movable part 22 located below the fixed part 21, and a spring member 23 connecting the fixed part 21 and the movable part 22 and resisting downward movement of the movable part 22 relative to the fixed part 21, and a movable member 30 that engages with a support body 40 that supports the stored items and moves vertically together with the movable part 22, the fixed part 21 can protrude from an upper end 30a of the movable member 30, and the amount of protrusion of the fixed part 21 relative to the upper end 30a of the movable member 30 changes as the vertical position of the movable member 30 changes depending on the load of the stored items.
Owner:DAIWA HOUSE INDUSTRY CO LTD

Location based audio signal message processing

PendingUS20260189872A1Display deviceMessage processing
A system is described that includes a display configured to visually display a mixed visual signal and audio signal, which includes a camera, a first microphone, a second microphone, two speakers, memory, a processor, that displays a mixed visual signal and the speakers emit a mixed audio signal.
Owner:ST VRTECH LLC

Multi-modal fusion oral English fluency real-time evaluation method and system

The invention provides a multi-modal fusion spoken English fluency real-time evaluation method and system, and relates to the technical field of intelligent education, the method comprises the following steps: collecting a multi-modal synchronous collection stream of a learner, the multi-modal synchronous collection stream comprising voice modal data, visual modal data and interaction modal data; extracting face key geometric feature points from the visual modal data, and calculating spatial correction parameters based on the face key geometric feature points; performing geometric correction on the visual modal data by using the spatial correction parameters, and performing time synchronization with the voice modal data to obtain visual features; performing acoustic feature extraction on the voice modal data to obtain voice features; based on time sequence information in the interaction modal data, extracting interaction features; and obtaining multi-modal feature data based on the visual features, the voice features and the interaction features. According to the invention, a closed loop from multi-mode synchronous acquisition and fusion analysis to real-time feedback is realized.
Owner:HENAN POLICE ACAD

Method and control unit for maintaining a passenger transport system

The invention relates to a method for maintaining a passenger transport system (14). The method comprises: receiving image data (BD) of a scene (12) having a component (16) of a passenger transport system (14); determining which component (16) of the passenger transport system (14) relates by means of the digital twin of the passenger transport system (14); determining at least one characteristic of the component (16) by means of twinning data (DD) of the digital twinning; generating a display signal (AZ) for a visual display (30) such that a characteristic of the component (16) is visually displayed in association therewith; comparing the component (16) represented by the image data (BD) with a digital representation of the component (16) by means of the image data (BD) and the twin data (DD); and determining whether the component (16) should be replaced.
Owner:INVENTIO AG

Selective visual display

According to one aspect of the invention, a multi-modal reading system is provided, comprising: a display screen; an eye movement tracking device; a processor; one or more computer storage devices; and a speaker, wherein the processor is configured to: measure, by the eye movement tracking device, eye movement of the user to determine a particular word gazed by the user; modifying the display of the text content, and visually highlighting the word or phrase of the current fixation point of the user; calculating the number of words read by the user per minute based on continuous measurements of the specific words watched by the user; according to the calculated reading rate, delay is set between continuous text element presentation so as to adapt to user cognition processing time; and playing the text content to the user at a speech speed matched with the reading speed of the user through a loudspeaker.
Owner:德查姆斯·理查德·克里斯托弗

System and method for determining liquid consumption by domestic animal from water dispenser

PendingUS20260248108A1Visual monitoringFishery
A method for monitoring and determining consumption from a water dispenser by an animal, the method comprising: a) an animal approaching the water dispenser and coming into a view of a visual monitoring device; b) the animal consuming (e.g.. lap. swallow) a liquid from a serving area of the water dispenser; c) the visual monitoring device receiving incoming video data of the animal consuming the liquid, wherein the incoming video data includes: i) at least a portion of a face and / or neck (e.g.. mouth, tongue, cheek, neck) of the animal visually indicating the animal is consuming the liquid; ii) the surface area of the liquid in serving area; iii) an inner surface of the serving area; and / or iv) a plurality of measurement aids in the serving area; d) automatically determining an amount of the liquid consumed by the animal based on the incoming video data.
Owner:AUTOMATED PET CARE PRODUCTS LLC

Precipitation identification method, device and equipment based on visual intelligent perception and storage medium

The invention provides a precipitation identification method, device and equipment based on visual intelligent perception, and a storage medium. The precipitation identification method based on visual intelligent perception comprises the following steps: obtaining a precipitation monitoring video stream, humidity data and temperature data; according to the rainfall monitoring video stream and the rainfall model, determining a predicted rainfall intensity type and a visual modal confidence coefficient corresponding to the predicted rainfall intensity type, and determining a predicted rainfall intensity type corresponding to accurate rainfall intensity based on the rainfall model; the rainfall comprehensive confidence coefficient is determined according to the humidity data, the temperature data and the visual modal confidence coefficient, the rainfall comprehensive confidence coefficient is determined by combining the humidity and the temperature on the visual basis, the defect of visual recognition can be overcome, and when the rainfall comprehensive confidence coefficient is larger than or equal to a decision confidence coefficient threshold value, the rainfall comprehensive confidence coefficient is determined. And the predicted rainfall intensity type is used as a rainfall identification result, so that the accuracy of rainfall identification is improved.
Owner:SHENZHEN NAT CLIMATE OBSERVATORY (SHENZHEN OBSERVATORY)

Multi-level emergency command intelligent decision method and system based on knowledge graph

PendingCN122155461AData processing applicationsBiological modelsEntity–relationship modelNetwork output
The application discloses a multi-level emergency command intelligent decision-making method and system based on a knowledge graph. The method comprises the following steps: S1, extracting the exclusive features of text modalities, visual modalities and space-time modalities from multi-source heterogeneous data in the whole scene of multi-level emergency command; S2, constructing a cross-modal attention fusion network to output multi-modal fusion features; S3, constructing a multi-modal joint entity relationship Transformer model, taking the multi-modal fusion features as input, and outputting standardized entity-relation triplets; S4, based on the entity-relation triplets, constructing a three-level hierarchical knowledge graph; S5, generating an emergency decision-making scheme by using a three-layer reasoning framework, and outputting the decision-making comprehensive confidence of the scheme; and S6, according to a preset emergency response level, combining the decision-making comprehensive confidence, and executing corresponding human-machine collaborative rules. In this way, intelligent sensing, accurate judgment, collaborative decision-making, efficient disposal and closed-loop optimization of the whole life cycle of an emergency event are realized.
Owner:BEIJING SMART SHANGQI TECHNOLOGY CO LTD

Intelligent monitoring system and control method thereof

The invention provides an intelligent monitoring system and a control method thereof. In the system, a laser radar detection assembly (1) is used for detecting actual position coordinates of a target in a conventional environment and a single visual mode in a complex environment of low illumination, rain and fog; the optical detection assembly (2) enables a visible light image and an infrared image of a detected intruding target to coincide, so that optical information feature data of visible light and infrared light can be accurately fused; the control assembly (3) is used for accurately determining the category attribute and coordinate information of a target object according to the actual position coordinates of the target in different environments and the accurately fused optical information feature data, displaying the identification information of the intruding target on the display screen (6), and controlling the on-off state of the emission trigger device (7) according to the identification information. The alarm is controlled to execute, and the locking or transmitting state of the transmitting device (8) is controlled. The method is high in recognition precision and wide in application range, and has active early warning capability and high response efficiency.
Owner:HEILONGJIANG NORTH TOOLS CO LTD

system

We provide the system. [Solution] A means of acquiring audio for collection, A processing means for pre-processing the acquired audio and extracting features, An analysis means for classifying speech based on extracted features and determining appropriate suggestions, A means of displaying the decided proposal to the user, A communication means for physically or visually transmitting voice analysis results using human assistance devices, A system that includes this.
Owner:SOFTBANK GROUP CORP

origami

This origami design replaces the traditional, language-based instruction of origami folding, which focuses on directional instructions such as up, down, left, and right placement, rotation, and fold lines. Instead, it allows learners, regardless of age, language, or experience, to easily and intuitively understand the placement and direction of origami through visual means, thus functioning as a universal design. [Solution] This origami is characterized by having at least three identification marks 2 of different shapes drawn on each of the four corners of a rectangular piece of paper 1. Preferably, at least two of the identification marks are not point-symmetrical, the identification marks are drawn in black, and there are four identification marks in total. Even more preferably, a mark 3 indicating the center position of the rectangular paper is drawn.