Urban scene image-text alignment method and system based on spatial topology constraint

Through the Transformer recoding and spatial topological constraint methods, the problem of insufficient modal information fusion in urban scene picture and text alignment is solved, and the consistent alignment of image and text features is achieved, which improves the accuracy of alignment tasks and the cognitive ability of the model.

CN120339761APending Publication Date: 2025-07-18CHINA UNIV OF MINING & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510148787.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing urban scene picture and text alignment methods cannot fully explore the potential connection between images and text, resulting in insufficient fusion of multimodal information and affecting task performance.

Method used

Transformer is used to recode different modal data, use spatial topological constraints to describe the relationship and constraints between image and text features, minimize the differences between modalities, and complete the graphic and text alignment task through the CLIP alignment module.

Benefits of technology

The geometric consistency and relational consistency of image and text features are achieved, and the accuracy of urban scene picture and text alignment and the model's cognitive ability of complex scenes are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339761A_ABST
    Figure CN120339761A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of multi-modal image-text alignment, and particularly relates to an urban scene image-text alignment method and system based on spatial topology constraint, and the method comprises the following steps: firstly, respectively extracting corresponding features of text data and image data by utilizing Transform, and encoding, so as to solve the problem of representation difference between different modal data types; the method comprises the following steps of: firstly, carrying out image-text alignment on an image, then describing a relationship and constraint between the image and text features by utilizing spatial topology constraint, minimizing differences among different modes, keeping geometric consistency and relationship consistency among entities, and finally, inputting the constrained features into a CLIP alignment module to finish an image-text alignment task. According to the method, by fusing image and text information, the cognitive ability of the model to a complex scene is enhanced, and the method has an important promotion effect in the fields of multi-modal data processing and deep learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of multimodal image-text alignment, and particularly relates to a method and system for urban scene image-text alignment based on spatial topological constraints. Background Technique

[0002] With the development of computer vision and natural language processing technologies, multimodal learning of images and texts has gradually become a popular research field. Especially in urban environments, the relationship between images and texts is very complex, involving the interaction of various text and visual elements such as urban buildings, signs, traffic signs, and street names. The urban scene image-text alignment task involves associating visual elements in images with text information to achieve deeper understanding and analysis. Urban scene image-text alignment aims to enhance the understanding and analysis ability of urban environments by understanding the relationship between urban scenes in images and relevant text descriptions.

[0003] The research on image-text alignment can not only help understand different elements in urban environments, but also provide important support for tasks such as autonomous driving, intelligent city management, and information extraction. In addition, the urban scene image-text alignment task can also promote the development of cross-modal learning. By integrating image and text information, it enhances the model's cognitive ability for complex scenes, especially playing an important promoting role in the fields of multimodal data processing and deep learning.

[0004] Although the urban scene image-text alignment task has made certain progress in research and applications, it still faces many challenges and problems, including the differences in semantic expressions and feature representations between different modalities of urban scene data; there are often model technology differences when processing different text and image data; and there are also difficulties in the understanding and fusion of cross-modal information. Existing image-text alignment methods cannot fully explore the potential connection between images and texts, resulting in insufficient fusion of multimodal information, which in turn affects the performance of the task. Summary of the Invention

[0005] In view of the problems existing in the prior art, the present invention provides a method and system for urban scene image-text alignment based on spatial topological constraints, which can fully explore the potential connection between images and texts, fully integrate multimodal information, achieve cross-modal maintenance of geometric consistency and relationship consistency between entities, and complete image-text alignment.

[0006] To solve the above technical problems, the present invention provides the following technical solution: A method for urban scene image-text alignment based on spatial topological constraints, comprising the following steps:

[0007] S1. Use a Transformer to separately extract corresponding features from different modality data and perform re-encoding, and then map each modality feature to a shared space; the modality data includes text data and image data;

[0008] S2. Use spatial topological constraints to describe the relationships and constraints between different modal features, minimize the differences between different modalities, and maintain geometric consistency and relationship consistency between entities.

[0009] S3. Input the constrained features into the CLIP alignment module to complete the image-text alignment task.

[0010] Further, in the aforementioned step S1, for the text data {w1, w2,... w i ..., w n}, extract the corresponding features and perform re-encoding, and then map the text features into a shared space, including the following sub-steps:

[0011] A-1.1. The Transformer uses the self-attention mechanism to process the input sequence, performs word embedding on the text data {w1, w2,... w i ..., w n} to convert it into vectors of a fixed dimension, each word is mapped into a vector space, and the embedding representation of the entire input sequence is a matrix, as follows:

[0012]

[0013] Among them, d is the fixed dimension of the word vector, w i is the i-th word;

[0014] A-1.2. Add positional encoding to the word vectors to ensure that the representation of each word depends not only on the word embedding but also on its position in the sequence. The positional encoding PE is constructed using the sine function, as follows:

[0015]

[0016] Among them, pos is the position of the word, i is the dimension index, and d is the fixed dimension of the word vector;

[0017] The input word vector representation is:

[0018] X = W + PE T

[0019] A-1.3. The input data passes through several self-attention layers of the Transformer to calculate the relationships between the query matrix Q, the key matrix K, and the value matrix V. The calculation formula for each self-attention layer is:

[0020]

[0021] Among them, Q = XW Q , K = XWK , V = XW V , W Q , W K , W V are the learning weights, d k is the dimension of the key.

[0022] A-1.4. The output of the self-attention layer is connected to the original input X with a residual connection and layer normalization is performed to obtain X':

[0023] X' = LN(X + Att_output),

[0024] where Att_output is the output of the self-attention layer.

[0025] A-1.5. The input is fed into the feed-forward network to obtain the output features:

[0026] F T = LN(X' + ReLU(X'W1 + b1)W2 + b2)

[0027] where W1, W2 are the network weights, b1, b2 are the network biases; LN is layer normalization, and ReLU is the activation function.

[0028] Furthermore, in the aforementioned step S1, for the image data the corresponding features are extracted and re-encoded, and then the image features are mapped into a shared space, including the following sub-steps:

[0029] B-1.1. The image is divided into several blocks of size P×P, and the image is cut into blocks in total, and each block is flattened into a one-dimensional vector to be uniformly represented with the text, where C is the number of channels;

[0030] B-1.2. Each block is mapped into a high-dimensional vector space through a linear projection:

[0031] z i = W e x i + b e

[0032] where W e is the mapping matrix, b e is the bias term;

[0033] The embedded representation of the image is where d is the embedding dimension;

[0034] B-1.3. Adding positional encoding to the image patches to preserve the spatial relationship of the image patches in the image: Using a sequence to represent the position of the image patch in the entire image, the final image embedding is expressed as:

[0035] Y = Z + PE I

[0036] where, PE I(pos,i) = i, i = 1, 2,..., N;

[0037] B-1.4. Feeding the image data into the multi-head attention layer and processing it through the feed-forward network to obtain the output features:

[0038] Y' = LN(Y + Att_output)

[0039] F I = LN(Y' + ReLU(Y'W1 + b1)W2 + b2)

[0040] In the formula, W1, W2 are network weights, b1, b2 are network biases; LN is layer normalization, ReLU is the activation function, and Att_output is the output of the self-attention layer.

[0041] Furthermore, the aforementioned step S2 includes the following sub-steps:

[0042] S2.1. Constructing the spatial topology graph: Using the spatial graph attention mechanism for spatial topology constraint, constructing a spatial topology graph, that is, a feature adjacency matrix, for the input feature F = {f1, f2,....f i ..., f N}, as follows:

[0043]

[0044] where, α and β are weight coefficients used to balance the influence of spatial distance and feature similarity, and f i is a node feature; S2.2. Calculating the attention coefficient: Calculating the attention coefficient between nodes i and j, mapped through a learnable weight matrix, and adding a topological distance term to strengthen the spatial relationship:

[0045] e ij = LeakyRelu(a T [Wf i ||Wf j ) + λ·distance(i, j)

[0046] where, a is a trainable weight vector, || represents the feature concatenation operation, LeakyRelu is an activation function, and λ is a hyperparameter used to adjust the influence of spatial distance.

[0047] S2.3. Calculate the attention coefficients through the softmax operation:

[0048]

[0049] Among them, N(i) is the set of neighbors of node i, and it is set that those within a distance of one unit above, below, left, and right of the node are neighbor nodes; S2.4. Aggregate neighborhood information: Use the calculated attention coefficients to weight and aggregate the information of neighbor nodes:

[0050]

[0051] Among them, f i ′ is the updated feature of node i, and σ() is the activation function Relu;

[0052] S2.5. Minimize the modality difference: Use the method of contrastive learning to minimize the modality difference between image and text features and optimize the similarity between images and texts; For the text features after graph attention optimization and the image features define a contrastive loss function:

[0053]

[0054] Among them, P is the number of image-text pairs, and γ is a hyperparameter representing the maximum acceptable similarity gap between modalities.

[0055] Furthermore, the aforementioned step S3 includes the following sub-steps:

[0056] S3.1. Measure the quality of image-text alignment by calculating the similarity between the image representation and the text representation, and use cosine similarity to measure the similarity of these two embedding vectors:

[0057]

[0058] Among them, · represents the vector dot product, ||·|| represents the L2 norm of the vector, w′ i , z′ j are the text features and image features after spatial topological constraint respectively;

[0059] S3.2. The CLIP alignment module uses contrastive loss to optimize the module, and the loss function adopts the InfoNCE loss, as shown in the following formula:

[0060]

[0061] Among them, M is the batch size, and the loss function promotes the alignment of image-text pairs while maximizing the discrimination of negative samples.

[0062] S3.3. Finally, based on the cosine similarity of the text and the picture calculated by comparison, the text pair with the highest similarity is the one corresponding to the picture.

[0063] On the other hand, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of any method described in the present invention are implemented.

[0064] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any method described in the present invention are implemented.

[0065] Compared with the prior art, the beneficial technical effects of the present invention adopting the above technical solutions are as follows:

[0066] 1. A method for aligning urban scene text and pictures based on spatial topological constraints provided by the present invention can use Transformer re-encoding to unify data representation and solve the problem of representation differences between different data types.

[0067] 2. Design a multi-modal entity alignment method based on topological constraints to ensure the accuracy of the alignment result. Use spatial topological constraints to describe the relationship and constraints between image and text features, minimize the differences between different modalities, maintain the geometric consistency and relationship consistency between entities, and complete the task of aligning urban scene text and pictures. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 is a principle block diagram of a method for aligning urban scene text and pictures based on spatial topological constraints of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0069] In order to better understand the technical content of the present invention, specific embodiments are given and described in conjunction with the accompanying drawings as follows.

[0070] In the present invention, various aspects of the present invention are described with reference to the accompanying drawings, and many illustrative embodiments are shown in the drawings. The embodiments of the present invention are not limited to those described in the drawings. It should be understood that the present invention can be implemented by any of the various concepts and embodiments introduced above, as well as the concepts and embodiments described in detail below, because the concepts and embodiments disclosed in the present invention are not limited to any embodiment. In addition, some aspects disclosed in the present invention can be used alone, or in any suitable combination with other aspects disclosed in the present invention.

[0071] Refer to Figure 1 , the present invention provides a method for aligning urban scene text and pictures based on spatial topological constraints, including the following steps:

[0072] S1. Use Transformer to extract corresponding features from different modal data respectively, re-encode them, and then map each modal feature into a shared space; the modal data includes text data and image data;

[0073] S2. Use spatial topological constraints to describe the relationships and constraints between different modal features, minimize the differences between different modalities, and maintain the geometric consistency and relationship consistency between entities;

[0074] S3. Input the constrained features into the CLIP alignment module to complete the text-image alignment task.

[0075] As a preferred embodiment of the present invention, in step S1, for the text data {w1, w2,... w i ..., w n}, extract corresponding features, re-encode them, and then map the text features into a shared space, including the following sub-steps: A-1.1. Transformer uses the self-attention mechanism to process the input sequence, perform word embedding on the text data {w1, w2,... w i ..., w n} to convert it into vectors of a fixed dimension, each word is mapped into a vector space, and the embedding representation of the entire input sequence is a matrix, as follows:

[0076]

[0077] where, d is the fixed dimension of the word vector, and w i is the i-th word;

[0078] A-1.2. Due to the disorder of Transformer itself, position encoding needs to be added to the word vectors to ensure that the representation of each word depends not only on the word embedding. The position encoding PE is constructed using the sine function, as follows:

[0079]

[0080] where pos is the position of the word, i is the dimension index, and d is the fixed dimension of the word vector;

[0081] The input word vector representation is:

[0082] X = W + PE T

[0083] A-1.3. The input data passes through several self-attention layers of Transformer, and the relationships between the query matrix Q, the key matrix K, and the value matrix V are calculated. The calculation formula of each self-attention layer is:

[0084]

[0085] Among them, Q = XW Q , K = XW K , V = XW V , W Q , W K , W V are the learning weights respectively, and d k is the dimension of the key. A-1.4. The output of the self-attention layer is connected with the original input X by residual connection and layer normalization is performed to obtain X':

[0086] X' = LN(X + Att_output),

[0087] where Att_output is the output of the self-attention layer.

[0088] A-1.5. The input is fed into the feed-forward network to obtain the output features:

[0089] F T = LN(X' + ReLU(X'W1 + b1)W2 + b2)

[0090] where W1, W2 are the network weights, b1, b2 are the network biases; LN is layer normalization, and ReLU is the activation function.

[0091] As a preferred embodiment of the present invention, in step S1, for image data corresponding features are extracted and re-coded, and then the image features are mapped into a shared space, including the following sub-steps: B-1.1. The image is divided into several blocks of size P×P, and the image is cut into a total of blocks, and each block is flattened into a one-dimensional vector to be uniformly represented with the text, where C is the number of channels; in an RGB image, C = 3. B-1.2. Each block is mapped into a high-dimensional vector space through a linear projection:

[0092] z i = W e x i + b e

[0093] where W e is the mapping matrix, and b e is the bias term;

[0094] The embedded representation of the image is where d is the embedding dimension;

[0095] B-1.3. Add positional encoding to the image patches to preserve the spatial relationship of the image patches in the image: Use a sequence to represent the position of the image patch in the entire image. The final image embedding is expressed as:

[0096] Y = Z + PE I

[0097] where, PE I(pos,i) = i, i = 1, 2,..., N;

[0098] B-1.4. Feed the image data into the multi-head attention layer and process it through the feed-forward network to obtain the output features:

[0099] Y' = LN(Y + Att_output)

[0100] F I = LN(Y' + ReLU(Y'W1 + b1)W2 + b2)

[0101] In the formula, W1, W2 are network weights, b1, b2 are network biases; LN is layer normalization, ReLU is the activation function, and Att_output is the output of the self-attention layer

[0102] Further, step S2 includes the following sub-steps:

[0103] S2.1. Considering the importance of the spatial structure of the image data, through the use of the spatial graph attention mechanism for spatial topology constraint, the features extracted by the Transformer not only contain global information but also can model the relationship of local neighborhoods more finely in space, improving the adaptability to spatial constraints.

[0104] Construct a spatial topology graph: Use the spatial graph attention mechanism for spatial topology constraint. For the input feature F = {f1, f2,....f i ..., f N}, construct a spatial topology graph, that is, a feature adjacency matrix, as follows:

[0105]

[0106] where, α and β are weight coefficients used to balance the influence of spatial distance and feature similarity, and f i is a node feature; S2.2. Calculate the attention coefficient: Calculate the attention coefficient between nodes i and j, which is mapped by a learnable weight matrix, and at the same time add a topological distance term to strengthen the spatial relationship:

[0107] e ij = LeakyRelu(a T [Wf i ||Wf j) + λ·distance(i, j)

[0108] Among them, a is a trainable weight vector, || represents the feature concatenation operation, LeakyRelu is an activation function, and λ is a hyperparameter used to adjust the influence of spatial distance.

[0109] S2.3. Calculate the attention coefficient through the softmax operation:

[0110]

[0111] Among them, N(i) is the neighbor set of node i, and nodes within a distance of one unit above, below, left, and right of node i are set as neighbor nodes; S2.4. Aggregate neighborhood information: Use the calculated attention coefficient to weight and aggregate the information of neighbor nodes:

[0112]

[0113] Among them, f i ′ is the updated feature of node i, and σ() is the activation function Relu;

[0114] S2.5. Minimize the modality difference: Use the method of contrastive learning to minimize the modality difference between image and text features and optimize the similarity between image and text; For the text features after graph attention optimization and image features

[0115]

[0116] Define a contrastive loss function:

[0117] As a preferred embodiment of the present invention, step S3 includes the following sub-steps:

[0118] S3.1. Measure the quality of image-text alignment by calculating the similarity between the image representation and the text representation, and use cosine similarity to measure the similarity of these two embedding vectors:

[0119]

[0120] Among them, · represents the dot product of vectors, |||| represents the L2 norm of vectors, w′ i , z′ j are the text features and image features after spatial topological constraint respectively;

[0121] S3.2. CLIP uses contrastive loss to optimize its modules. The goal is to maximize the similarity between each pair of images and corresponding texts, while minimizing the similarity between different pairs. The loss function uses InfoNCE loss, as shown in the following equation:

[0122]

[0123] where M is the batch size. The loss function promotes the alignment of image-text pairs while maximizing the discrimination of negative samples.

[0124] S3.3. Finally, based on the cosine similarity between the text and the image calculated by comparison, the text pair with the highest similarity is the one corresponding to the image.

[0125] On the other hand, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of any one of the methods described in the present invention are implemented.

[0126] The present invention also provides a computer-readable storage medium, on which a computer program is stored. The computer program is characterized in that when it is executed by a processor, the steps of any one of the methods described in the present invention are implemented.

[0127] Although the present invention has been described above with reference to preferred embodiments, it is not intended to limit the present invention. Those with ordinary knowledge in the technical field to which the present invention pertains may make various modifications and refinements without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be determined by the scope defined in the claims.

Claims

1. An urban scene graphic-text alignment method based on spatial topological constraints, characterized in that, It includes the following steps: S1. Use a Transformer to extract corresponding features from different modal data respectively and perform re-encoding, and then map each modal feature into a shared space; the modal data includes text data and image data; S2. Use spatial topology constraints to describe the relationships and constraints between different modal features, minimize the differences between different modalities, and maintain the geometric consistency and relationship consistency between entities; S3. Input the constrained features into the CLIP alignment module to complete the image-text alignment task.

2. The method for aligning urban scene text and images based on spatial topology constraints according to claim 1, wherein In step S1, corresponding features are extracted from the text data {w1, w2,... w i ..., w n} and re-encoded, and then the text features are mapped into a shared space, including the following sub-steps: A-1.

1. The Transformer uses the self-attention mechanism to process the input sequence. For the text data {w1, w2,... w i ..., w n}, word embeddings are performed to convert it into vectors of a fixed dimension. Each word is mapped into a vector space, and the embedding representation of the entire input sequence is a matrix, as follows: Among them, d is the fixed dimension of the word vector, and w i is the i-th word; A-1.

2. Add positional encoding to the word vectors to ensure that the representation of each word depends not only on the word embedding but also on its position in the sequence. The positional encoding PE is constructed using a sine function as follows: where pos is the position of the word, i is the dimension index, and d is the fixed dimension of the word vector; The input word vector is represented as: X = W + PE T A-1.

3. The input data passes through several self-attention layers of the Transformer to calculate the relationships between the query matrix Q, the key matrix K, and the value matrix V. The calculation formula for each self-attention layer is: where Q = XW Q , K = XW K , V = XW V , W Q , W K , W V are the learning weights respectively, and d k is the dimension of the key. A-1.

4. The output of the self-attention layer is connected to the original input X by residual connection and layer normalization is performed to obtain X': X′ = LN(X + Att_output), where Att_output is the output of the self-attention layer. A-1.

5. Input the data into the feed-forward network to obtain the output features: F T = LN(X′ + ReLU(X′W1 + b1)W2 + b2) where W1, W2 are the network weights, b1, b2 are the network biases; LN is layer normalization, and ReLU is the activation function.

3. A method for aligning urban scene text and images based on spatial topological constraints according to claim 1, characterized in that In step S1, for the image data Extract the corresponding features and perform re-encoding, and then map the image features into a shared space, including the following sub-steps: B-1.

1. Divide the image into several blocks of size P×P. The image is cut into a total of blocks, and flatten each block into a one-dimensional vector for unified representation with the text, where C is the number of channels; B-1.

2. Map each block into a high-dimensional vector space through a linear projection: z i = W e x i + b e Among them, W e is the mapping matrix, and b e is the bias term; The embedded representation of the image is where d is the embedding dimension; B-1.

3. Add positional encoding to the image blocks to retain the spatial relationships of the image blocks in the image: Use sequences to represent the positions of the image blocks in the entire image. The final image embedding is represented as: Y = Z + PE I Among them, PE I(pos,i) = i, i = 1, 2,..., N; B-1.

4. Input the image data into the multi-head attention layer and process it through the feed-forward network to obtain the output features: Y′ = LN(Y + Att_output) F I = LN(Y′ + ReLU(Y′W1 + b1)W2 + b2) In the formula, W1, W2 are the network weights, b1, b2 are the network biases; LN is layer normalization, ReLU is the activation function, and Att_output is the output of the self-attention layer.

4. A method for aligning urban scene text and images based on spatial topology constraints according to claim 1, characterized in that, Step S2 includes the following sub-steps: S2.

1. Construct a spatial topology graph: Use the spatial graph attention mechanism for spatial topology constraints. For the input feature F = {f1, f2,....f i ..., f N}, construct a spatial topology graph, that is, a feature adjacency matrix, as follows: where α and β are weight coefficients used to balance the influence of spatial distance and feature similarity, and f i is a node feature; S2.

2. Calculate the attention coefficients: Calculate the attention coefficients between nodes i and j, which are mapped through a learnable weight matrix, and at the same time add a topological distance term to strengthen the spatial relationship: e ij = LeakyRelu(a T [Wf i ||Wf j ) + λ·distance(i, j) where a is a trainable weight vector, || represents the feature concatenation operation, LeakyRelu is an activation function, and λ is a hyperparameter used to adjust the influence of the spatial distance. S2.

3. Calculate the attention coefficients through the softmax operation: where N(i) is the set of neighbors of node i, and it is set that the nodes within a distance of one unit up, down, left, and right from node i are all neighbor nodes; S2.

4. Aggregate the neighborhood information: Use the calculated attention coefficients to weighted-aggregate the information of the neighbor nodes: Among them, f i ' is the updated feature of node i, and σ() is the activation function Relu; S2.

5. Minimize the modality difference: Using the method of contrastive learning, minimize the modality difference between image and text features, and optimize the similarity between images and texts; for the text features optimized by graph attention and image features Define a contrastive loss function: where P is the number of image-text pairs, and γ is a hyperparameter representing the maximum acceptable similarity gap between modalities.

5. A method for aligning urban scene text and images based on spatial topological constraints according to claim 1, characterized in that Step S3 includes the following sub-steps: S3.

1. Measure the quality of image-text alignment by calculating the similarity between the image representation and the text representation, and use cosine similarity to measure the similarity of these two embedding vectors: where, · represents the dot product of vectors, ||·|| represents the L2 norm of vectors, w′ i , z′ j are the text feature and the image feature after spatial topology constraint, respectively; S3.

2. CLIP uses contrastive loss to optimize its module, and the loss function adopts InfoNCE loss, as shown in the following formula: where M is the batch size. The loss function promotes the alignment of image-text pairs while maximizing the discrimination of negative samples. S3.

3. Finally, based on the calculated cosine similarity between the text and the image, the text pair with the highest similarity is the text pair corresponding to the image.

6. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 5.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method described in any one of claims 1 to 5.