Special Session 5. Global Frontiers in AI-Driven Intelligent Image/Video Processing and related Innovation

SUBMIT ONLINE: https://www.easychair.org/conferences/?conf=prai2026  (Please choose Special Session 5)

The field of AI-Driven Intelligent Image/Video Processing represents a transformative frontier in computer vision, where advanced machine learning models are revolutionizing how data is analyzed, enhanced, and synthesized. This domain encompasses innovations in real-time object detection, scene understanding, high-resolution image/video restoration, generative content creation, and multimodal fusion with inputs. Key applications span autonomous systems, healthcare, entertainment, surveillance, and creative industries. Recent breakthroughs include diffusion models for photorealistic image generation, transformer-based architectures for long-range video context modeling, and edge AI for low-latency processing. Challenges involve ethical considerations, computational efficiency, and robustness against adversarial attacks. This area drives cross-disciplinary collaboration, pushing boundaries in AI scalability, interpretability, and real-world deployment, while shaping future standards for intelligent media.

Chair:
Zhipan Wu, Huizhou University, China
Email: 19869841@qq.com

 

RELATED TOPICS
Topics of interest include, but are not limited to:

  • 1. AI-Powered Real-Time Low-Latency Image/Video Processing
    2. Intelligent Fields: AI-Driven Solutions for Agricultural Efficiency
    3. AI-Driven Innovations in Forestry and Ecosystem Conservation‌
    4. AI-Powered Medical Imaging: Diagnosis, Restoration & Synthetic Data
    5. Intelligent Applications for Elderly Care and Child Development
    6. Intelligent Edge Vision: AI-Powered Real-Time Monitoring and Analysis
    7. Smart Eyes on the Edge: AI-Driven Visual Processing for Autonomous Systems
    8. The Safety Frontier of AI Image/Video Processing
    9. ……

Special Session 5 - Invited Speakers

Kazuya Ueki, Meisei University, Japan
He received his B.S. degree in Information Engineering in 1997 and his M.S. degree in Computer and Mathematical Sciences in 1999, both from Tohoku University, Sendai, Japan. In 1999, he joined NEC Soft, Ltd., Tokyo, Japan, where he was mainly engaged in research on face recognition. He received his Ph.D. degree from the Graduate School of Science and Engineering, Waseda University, Tokyo, Japan, in 2007. From 2013 to 2017, he served as an Assistant Professor at Waseda University. He is currently an Associate Professor in the School of Information Science, Meisei University. His research interests include information retrieval, video anomaly detection, pattern recognition, and machine learning. He is involved in the video retrieval evaluation benchmark (TRECVID) sponsored by the National Institute of Standards and Technology (NIST), contributing to the development of video retrieval technology. His submitted systems achieved the highest performance in the TRECVID AVS task in 2016, 2017, 2022, and 2025.
Speech Title: Video Retrieval over the Years: Lessons from a Decade of the Ad-hoc Video Search Task in the TRECVID Benchmark
Abstract: Over the past decade, the Ad-hoc Video Search (AVS) task in the TRECVID benchmark has driven the evolution of video retrieval. Early methods relied on predefined visual concepts. CLIP-style embeddings then enabled direct text-video matching in a shared semantic space. Today, vision-language models (VLMs) reframe retrieval as a broader task of multimodal understanding. They interpret user intent, reason about visual content, and support interactive exploration. This talk traces that trajectory through several approaches developed along the way. Query expansion using large language models and image generation enriches user queries and improves retrieval robustness. Caption-driven adaptation customizes retrieval models to new domains without manual annotation. VLM-based semantic verification uses visual question answering to re-rank results and better align them with user intent. The talk also shares experimental findings and practical experiences gained through years of participation in TRECVID, along with challenges that remain for future work.