•  
  •  
 

Abstract

The increasing volume of videos produced and shared via smart devices has made indexing and retrieving them a clear challenge. Current retrieval systems, such as specialized video retrieval systems, rely on text-based searches, which often misrepresent and frequently result in erroneous descriptions of the requested video, particularly when it comes to videos about certain specialities like security or medical procedures that are hard to describe. As a result, we developed a model that handles information-rich video queries, making it easier to find related videos that are specifically relevant to the user's needs. This paper proposes a new method for indexing and retrieving videos, aiming to preserve small-sized information while being a powerful and fast retrieval tool. The process involves three stages: preprocessing and indexing, which involves preparing each video with three basic keys; extracting features using a hybrid Transformer-based encoder with multiple algorithms and neural networks; and storing them in a lightweight CSV file. The final stage utilizes the modified RSA Reptile Search Algorithm to retrieve videos efficiently, achieving 1.0 at Accuracy@10, Precision@10, and mAP, while Recall remains a challenge due to the nature of the data. This approach requires lightweight resources and efficient modifications to handle the massive number of video files.

Keywords

Multimedia information retrieval, Reptile search algorithm, Similarity measures, Video retrieval, Vision transformer

Subject Area

Computer Science

Article Type

Article

First Page

2996

Last Page

3009

Creative Commons License

Creative Commons Attribution 4.0 International License
This work is licensed under a Creative Commons Attribution 4.0 International License.

Share

 
COinS