ביטויים רגולריים (Regex)

מה זה Regex?

ביטויים רגולריים (Regular Expressions או בקיצור Regex) הם תבניות המשמשות לחיפוש, התאמה ומניפולציה של טקסט. הם כלי רב-עוצמה לעבודה עם מחרוזות.

┌─────────────────────────────────────────────────┐
│                    Regex                        │
│  ┌──────────────┐    ┌──────────────────────┐   │
│  │  תבנית חיפוש  │ → │  התאמה במחרוזת       │   │
│  │   /hello/    │    │  "hello world"       │   │
│  └──────────────┘    └──────────────────────┘   │
└─────────────────────────────────────────────────┘

יצירת Regex ב-JavaScript

שתי דרכים ליצירת Regex

// 1. Literal notation - הדרך הנפוצה
const regex1 = /hello/;

// 2. Constructor - שימושי כשהתבנית דינמית
const regex2 = new RegExp('hello');

// עם דגלים (flags)
const regex3 = /hello/gi;
const regex4 = new RegExp('hello', 'gi');

דגלים (Flags)

דגלים משנים את התנהגות החיפוש:

דגל שם משמעות
g global מוצא את כל ההתאמות, לא רק הראשונה
i case-insensitive התעלם מהבדלי אותיות גדולות/קטנות
m multiline התייחס לכל שורה כמחרוזת נפרדת
s dotAll נקודה תתאים גם לשורה חדשה
u unicode תמיכה מלאה ב-Unicode
// חיפוש ללא רגישות לאותיות
const regex = /hello/i;
console.log(regex.test('HELLO'));  // true
console.log(regex.test('Hello'));  // true

// מציאת כל ההתאמות
const str = 'cat cat cat';
const matches = str.match(/cat/g);
console.log(matches);  // ['cat', 'cat', 'cat']

תווים מיוחדים

Meta Characters - תווי מטא

תו משמעות דוגמה
. כל תו בודד (חוץ משורה חדשה) a.c מתאים ל-abc, aXc
\d ספרה (0-9) \d\d מתאים ל-42
\D לא ספרה \D מתאים ל-a, !
\w תו מילה (a-z, A-Z, 0-9, _) \w+ מתאים ל-hello_123
\W לא תו מילה \W מתאים ל-@,
\s רווח לבן (space, tab, newline) \s+ מתאים לרווחים
\S לא רווח לבן \S+ מתאים למילה
\b גבול מילה \bcat\b מתאים ל-cat אך לא ל-category
\n שורה חדשה
\t טאב
// חיפוש ספרות
const phone = '054-1234567';
console.log(phone.match(/\d+/g));  // ['054', '1234567']

// חיפוש מילים
const text = 'Hello World 123';
console.log(text.match(/\w+/g));  // ['Hello', 'World', '123']

כימות (Quantifiers)

כמה פעמים התו יופיע

כימות משמעות דוגמה
* 0 או יותר ab*c → ac, abc, abbc
+ 1 או יותר ab+c → abc, abbc (לא ac)
? 0 או 1 colou?r → color, colour
{n} בדיוק n פעמים a{3} → aaa
{n,} n או יותר a{2,} → aa, aaa, aaaa...
{n,m} בין n ל-m פעמים a{2,4} → aa, aaa, aaaa
// מספר טלפון ישראלי
const phoneRegex = /^0\d{1,2}-?\d{7}$/;
console.log(phoneRegex.test('054-1234567'));  // true
console.log(phoneRegex.test('02-1234567'));   // true
console.log(phoneRegex.test('0541234567'));   // true

// כתובת דוא"ל פשוטה
const emailRegex = /\w+@\w+\.\w+/;
console.log(emailRegex.test('[email protected]'));  // true

קבוצות וטווחים

Character Classes - מחלקות תווים

// טווח של אותיות
const regex1 = /[a-z]/;     // אות קטנה אחת
const regex2 = /[A-Z]/;     // אות גדולה אחת
const regex3 = /[0-9]/;     // ספרה (כמו \d)
const regex4 = /[a-zA-Z]/;  // כל אות

// תווים ספציפיים
const vowels = /[aeiou]/gi;
const text = 'Hello World';
console.log(text.match(vowels));  // ['e', 'o', 'o']

// שלילה עם ^
const notDigit = /[^0-9]/;  // כל תו שהוא לא ספרה
console.log('a1b2'.match(/[^0-9]/g));  // ['a', 'b']

Groups - קבוצות

// קבוצות לכידה (Capturing Groups)
const regex = /(\d{2})-(\d{2})-(\d{4})/;
const date = '25-12-2024';
const match = date.match(regex);

console.log(match[0]);  // '25-12-2024' (כל ההתאמה)
console.log(match[1]);  // '25' (קבוצה ראשונה)
console.log(match[2]);  // '12' (קבוצה שנייה)
console.log(match[3]);  // '2024' (קבוצה שלישית)

// קבוצות עם שמות (Named Groups)
const namedRegex = /(?<day>\d{2})-(?<month>\d{2})-(?<year>\d{4})/;
const namedMatch = date.match(namedRegex);

console.log(namedMatch.groups.day);    // '25'
console.log(namedMatch.groups.month);  // '12'
console.log(namedMatch.groups.year);   // '2024'

עוגנים (Anchors)

עוגנים מציינים מיקום במחרוזת:

עוגן משמעות
^ תחילת המחרוזת
$ סוף המחרוזת
\b גבול מילה
\B לא גבול מילה
// מתחיל עם "Hello"
const startsWithHello = /^Hello/;
console.log(startsWithHello.test('Hello World'));  // true
console.log(startsWithHello.test('Say Hello'));    // false

// מסתיים עם ספרות
const endsWithNumber = /\d+$/;
console.log(endsWithNumber.test('version 2'));    // true
console.log(endsWithNumber.test('2 items'));      // false

// מילה שלמה בלבד
const wholeWord = /\bcat\b/;
console.log(wholeWord.test('cat'));        // true
console.log(wholeWord.test('category'));   // false
console.log(wholeWord.test('a cat here')); // true

מתודות של Regex

test() - בדיקת התאמה

const regex = /hello/i;

// מחזיר true או false
console.log(regex.test('Hello World'));  // true
console.log(regex.test('Goodbye'));      // false

exec() - ביצוע חיפוש מפורט

const regex = /(\w+)@(\w+)\.(\w+)/;
const email = '[email protected]';

const result = regex.exec(email);
console.log(result[0]);  // '[email protected]'
console.log(result[1]);  // 'user'
console.log(result[2]);  // 'example'
console.log(result[3]);  // 'com'
console.log(result.index);  // 0 (מיקום ההתאמה)

מתודות מחרוזת עם Regex

match() - מציאת התאמות

const text = 'The rain in Spain stays mainly in the plain';

// התאמה ראשונה בלבד
console.log(text.match(/ain/));
// ['ain', index: 5, input: '...']

// כל ההתאמות עם g flag
console.log(text.match(/ain/g));
// ['ain', 'ain', 'ain', 'ain']

matchAll() - כל ההתאמות עם פרטים

const text = 'cat bat rat';
const regex = /[cbr]at/g;

for (const match of text.matchAll(regex)) {
    console.log(`נמצא: ${match[0]} במיקום ${match.index}`);
}
// נמצא: cat במיקום 0
// נמצא: bat במיקום 4
// נמצא: rat במיקום 8

replace() - החלפת טקסט

// החלפה פשוטה
const text = 'Hello World';
console.log(text.replace(/World/, 'JavaScript'));
// 'Hello JavaScript'

// החלפת כל ההתאמות
const text2 = 'cat cat cat';
console.log(text2.replace(/cat/g, 'dog'));
// 'dog dog dog'

// שימוש בקבוצות
const name = 'John Smith';
console.log(name.replace(/(\w+) (\w+)/, '$2, $1'));
// 'Smith, John'

// פונקציית החלפה
const prices = 'Items cost $5 and $10';
const converted = prices.replace(/\$(\d+)/g, (match, amount) => {
    return `₪${amount * 3.5}`;
});
console.log(converted);
// 'Items cost ₪17.5 and ₪35'

replaceAll() - החלפת כל ההתאמות

const text = 'a-b-c-d';
console.log(text.replaceAll('-', '_'));
// 'a_b_c_d'

// עם regex (חייב דגל g)
console.log(text.replaceAll(/-/g, '_'));
// 'a_b_c_d'

search() - חיפוש אינדקס

const text = 'Hello World';

console.log(text.search(/World/));  // 6
console.log(text.search(/xyz/));    // -1 (לא נמצא)

split() - פיצול מחרוזת

// פיצול לפי תבנית
const text = 'apple, banana; cherry  orange';
const fruits = text.split(/[,;\s]+/);
console.log(fruits);
// ['apple', 'banana', 'cherry', 'orange']

// פיצול עם שמירת המפריד
const text2 = 'one1two2three3';
const parts = text2.split(/(\d)/);
console.log(parts);
// ['one', '1', 'two', '2', 'three', '3', '']

דוגמאות שימושיות

ולידציה של נתונים

// ולידציית אימייל
function isValidEmail(email) {
    const regex = /^[^\s@]+@[^\s@]+\.[^\s@]+$/;
    return regex.test(email);
}

console.log(isValidEmail('[email protected]'));  // true
console.log(isValidEmail('invalid.email'));     // false

// ולידציית מספר טלפון ישראלי
function isValidPhoneIL(phone) {
    const regex = /^0(5[0-9]|[2-4]|[7-9])\d{7}$/;
    return regex.test(phone.replace(/-/g, ''));
}

console.log(isValidPhoneIL('054-1234567'));  // true
console.log(isValidPhoneIL('02-1234567'));   // true

// ולידציית סיסמה
function isStrongPassword(password) {
    // לפחות 8 תווים, אות גדולה, אות קטנה, ספרה
    const regex = /^(?=.*[a-z])(?=.*[A-Z])(?=.*\d).{8,}$/;
    return regex.test(password);
}

console.log(isStrongPassword('MyPass123'));  // true
console.log(isStrongPassword('weak'));       // false

// ולידציית ת.ז. ישראלית (9 ספרות)
function isValidIsraeliID(id) {
    const regex = /^\d{9}$/;
    return regex.test(id);
}

חילוץ מידע

// חילוץ כתובות URL
function extractURLs(text) {
    const regex = /https?:\/\/[^\s]+/g;
    return text.match(regex) || [];
}

const text = 'Visit https://example.com or http://test.org';
console.log(extractURLs(text));
// ['https://example.com', 'http://test.org']

// חילוץ האשטאגים
function extractHashtags(text) {
    const regex = /#\w+/g;
    return text.match(regex) || [];
}

const tweet = 'Love #JavaScript and #coding! #webdev';
console.log(extractHashtags(tweet));
// ['#JavaScript', '#coding', '#webdev']

// חילוץ מספרים
function extractNumbers(text) {
    const regex = /-?\d+\.?\d*/g;
    return text.match(regex)?.map(Number) || [];
}

console.log(extractNumbers('Price: 25.99, Discount: -5'));
// [25.99, -5]

ניקוי וטרנספורמציה

// הסרת HTML tags
function stripHTML(html) {
    return html.replace(/<[^>]*>/g, '');
}

console.log(stripHTML('<p>Hello <b>World</b></p>'));
// 'Hello World'

// המרה ל-camelCase
function toCamelCase(str) {
    return str.replace(/[-_\s]+(.)/g, (_, char) => char.toUpperCase());
}

console.log(toCamelCase('hello-world'));      // 'helloWorld'
console.log(toCamelCase('my_variable_name')); // 'myVariableName'

// הוספת פסיקים למספרים גדולים
function formatNumber(num) {
    return num.toString().replace(/\B(?=(\d{3})+(?!\d))/g, ',');
}

console.log(formatNumber(1234567));  // '1,234,567'

// ניקוי רווחים מיותרים
function cleanWhitespace(text) {
    return text.replace(/\s+/g, ' ').trim();
}

console.log(cleanWhitespace('  Hello    World  '));
// 'Hello World'

Lookahead ו-Lookbehind

Lookahead - הסתכלות קדימה

// Positive Lookahead: (?=...)
// מצא "foo" שאחריו יש "bar"
const regex1 = /foo(?=bar)/;
console.log(regex1.test('foobar'));  // true
console.log(regex1.test('foobaz'));  // false

// Negative Lookahead: (?!...)
// מצא "foo" שאחריו אין "bar"
const regex2 = /foo(?!bar)/;
console.log(regex2.test('foobaz'));  // true
console.log(regex2.test('foobar'));  // false

Lookbehind - הסתכלות אחורה

// Positive Lookbehind: (?<=...)
// מצא מספר שלפניו יש $
const regex3 = /(?<=\$)\d+/;
console.log('$100'.match(regex3));  // ['100']

// Negative Lookbehind: (?<!...)
// מצא מספר שלפניו אין $
const regex4 = /(?<!\$)\d+/;
console.log('100'.match(regex4));   // ['100']
console.log('$100'.match(regex4));  // ['00'] (כי 1 בא אחרי $)

טעויות נפוצות

1. שכחתם דגל g

// לא נכון - מחליף רק את הראשון
const text = 'cat cat cat';
console.log(text.replace(/cat/, 'dog'));  // 'dog cat cat'

// נכון - מחליף את כולם
console.log(text.replace(/cat/g, 'dog')); // 'dog dog dog'

2. לא בריחה מתווים מיוחדים

// לא נכון - הנקודה זה wildcard
const regex1 = /example.com/;
console.log(regex1.test('exampleXcom'));  // true (לא נכון!)

// נכון - בריחה עם \
const regex2 = /example\.com/;
console.log(regex2.test('exampleXcom'));  // false
console.log(regex2.test('example.com'));  // true

3. Greedy vs Lazy

// Greedy (ברירת מחדל) - לוקח הכי הרבה
const text = '<div>Hello</div><div>World</div>';
console.log(text.match(/<div>.*<\/div>/)[0]);
// '<div>Hello</div><div>World</div>' (הכל!)

// Lazy (עם ?) - לוקח הכי מעט
console.log(text.match(/<div>.*?<\/div>/)[0]);
// '<div>Hello</div>' (רק הראשון)

סיכום

  • Regex הם כלי עוצמתי לעבודה עם טקסט
  • דגלים כמו g, i, m משנים התנהגות
  • תווי מטא כמו \d, \w, \s מייצגים קבוצות תווים
  • כימות כמו *, +, {n,m} קובעים כמה פעמים
  • קבוצות () מאפשרות לכידה והתייחסות
  • מתודות כמו test(), match(), replace() מפעילות regex

💡 טיפ: השתמשו בכלים אונליין כמו regex101.com לבדיקה ולימוד של ביטויים רגולריים!

בדקו את עצמכם

נסו לענות לבד לפני שאתם פותחים את התשובה.

  1. מה עושה הדגל g בביטוי רגולרי?

    1. מוצא את כל ההתאמות, לא רק את הראשונה
    2. עובד על כמה שורות
    3. שום דבר
    4. מתעלם מאותיות גדולות
    הצגת התשובה

    תשובה א. g הוא global. i מתעלם מאותיות גדולות, m לכמה שורות.

  2. מה מתאים ל-\d{3}?

    1. הספרה 3
    2. כל תו שלוש פעמים
    3. שלוש ספרות ברצף
    4. שלוש אותיות
    הצגת התשובה

    תשובה ג. \d היא ספרה, ו-{3} היא בדיוק שלוש פעמים.